feat(import): parse a PDF resume without an AI provider (#3400)

* feat(import): parse a PDF resume without an AI provider

Importing a PDF required a connected AI provider, so anyone without a
paid API key could only import the three JSON formats. Almost nobody
arrives with one of those files; they arrive with a PDF. The first thing
a new user tries to do was blocked behind bringing their own key.

Adds a deterministic parser that reads the text out of the PDF in the
browser and prefills the builder. It pulls the contact block, segments
the body on conventional headings, and maps entries to real items,
reusing the ATS period parser for dates so a date range is not mistaken
for a phone number.

Nothing is thrown away: header parts that do not map to a field go into
the description, and unrecognized headings become custom sections. The
imported sections are placed on the page so the result renders straight
away. Output is validated against the resume schema before it is
returned.

Text extraction groups items by baseline rather than trusting hasEOL,
and turns wide column gaps into a double space, which is what lets a
row split into company, position and location.

The AI path still runs when a provider is connected. Word import is
unchanged and still requires one.

Closes #3334

* fix(import): keep every section and entry the PDF actually contains

Review found three ways the parser lost or mangled content, all of them
reproducible.

A document whose first heading was not one of the known aliases never
started a section, because unknown-heading detection was gated on a
section already being open. Everything after it was swallowed as contact
header text. The header block is now bounded by where the contact
details stop, so a heading is recognized wherever it appears.

An entry spreading company, position and dates over three lines was
imported as two malformed items. A line that introduces an entry now
merges into the open entry instead of starting a second one.

An uppercase company such as ACME CORPORATION was read as a section
heading and fragmented the entry. A heading candidate followed by a date
line is now treated as an entry header, which is what it is.

Also escape single quotes, and construct the PDF worker inside the try
so the nested worker is terminated even if construction throws.

Title-case headings are deliberately still not treated as headings:
company and school names are title case too, and splitting on them would
fragment real entries. Such a section stays in the preceding one with its
text intact rather than risking loss.

* fix(import): look past a multi-line preamble before calling a line a heading

The previous guard only inspected the next line, so an uppercase company
followed by a separate role line and then the dates was still read as a
section heading. The experience or education entry was moved into a
custom section and lost.

Heading detection now scans a two-line window for the date that marks an
entry, and stops early at a bullet so a genuine heading whose section
opens with bullet points is still recognized.

The window can suppress a real heading whose first entry puts a bare date
two lines below it. That is the deliberate direction to fail in: a missed
heading leaves the text in the preceding section, while a misread entry
fragments structured content.

* fix(import): collect an entry preamble until its dates appear

An entry that spread company, role, location and dates over four lines
was imported as two broken items: the company with no dates, and the
location carrying the period.

The cause was in entry grouping rather than heading detection. Lines
before a date were only folded into the entry header when the date sat
on the very next line; anything earlier fell through to the description.
Preamble lines are now collected into the entry header until the dates
turn up, bounded by the same lookahead and stopping at a bullet, so an
undated section cannot swallow itself.

The heading lookahead widens to four lines to match, which is the
realistic maximum for company, role, location and dates.

* fix(import): harden local PDF resume parsing

* chore(import): document audited HTML construction

---------

Co-authored-by: Amruth Pillai <im.amruth@gmail.com>
This commit is contained in:
Syed Ali Abbas Zaidi
2026-09-05 09:53:37 -07:00
committed by GitHub
co-authored by Amruth Pillai
parent ef47baf243
commit cce6d64afa
8 changed files with 1199 additions and 26 deletions
@@ -88,25 +88,54 @@ const renderDialog = () => {
);
};
// A real "%PDF" header so the dialog auto-detects the type and shows the provider notice.
const createPdfFile = () =>
new File([new Uint8Array([0x25, 0x50, 0x44, 0x46])], "resume.pdf", { type: "application/pdf" });
// Word still needs a provider, so it is what surfaces the notice. PDF falls back to a local parse.
const createWordFile = () =>
new File([new Uint8Array([0x50, 0x4b, 0x03, 0x04])], "resume.docx", {
type: "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
});
// The dialog renders through a portal, so query the document rather than the render container.
async function selectPdfFile() {
async function selectWordFile() {
const input = document.querySelector<HTMLInputElement>('input[type="file"]');
if (!input) throw new Error("File input not found");
fireEvent.change(input, { target: { files: [createPdfFile()] } });
fireEvent.change(input, { target: { files: [createWordFile()] } });
return await screen.findByText("Set up a provider");
}
const createPdfFile = () =>
new File([new Uint8Array([0x25, 0x50, 0x44, 0x46])], "resume.pdf", { type: "application/pdf" });
describe("ImportResumeDialog — PDF without a provider", () => {
it("offers a local parse instead of demanding an AI provider", async () => {
renderDialog();
const input = document.querySelector<HTMLInputElement>('input[type="file"]');
if (!input) throw new Error("File input not found");
fireEvent.change(input, { target: { files: [createPdfFile()] } });
expect(await screen.findByText(/read the text out of the PDF here in your browser/)).toBeInTheDocument();
expect(screen.queryByText("Set up a provider")).not.toBeInTheDocument();
});
it("keeps the import button usable", async () => {
renderDialog();
const input = document.querySelector<HTMLInputElement>('input[type="file"]');
if (!input) throw new Error("File input not found");
fireEvent.change(input, { target: { files: [createPdfFile()] } });
await screen.findByText(/read the text out of the PDF here in your browser/);
expect(screen.getByRole("button", { name: "Import" })).not.toBeDisabled();
});
});
describe("ImportResumeDialog — Set up a provider", () => {
// https://github.com/amruthpillai/reactive-resume/issues/3307
it("confirms before leaving instead of navigating behind the dialog", async () => {
renderDialog();
const link = await selectPdfFile();
const link = await selectWordFile();
fireEvent.click(link);
@@ -118,7 +147,7 @@ describe("ImportResumeDialog — Set up a provider", () => {
it("stays put and keeps the selected file when the user cancels", async () => {
renderDialog();
const link = await selectPdfFile();
const link = await selectWordFile();
fireEvent.click(link);
fireEvent.click(await screen.findByText("Stay"));
@@ -130,12 +159,12 @@ describe("ImportResumeDialog — Set up a provider", () => {
expect(routerNavigate).not.toHaveBeenCalled();
expect(navigate).not.toHaveBeenCalled();
expect(useDialogStore.getState().open).toBe(true);
expect(screen.getByText("resume.pdf")).toBeInTheDocument();
expect(screen.getByText("resume.docx")).toBeInTheDocument();
});
it("closes the dialog and navigates once the user confirms", async () => {
renderDialog();
const link = await selectPdfFile();
const link = await selectWordFile();
fireEvent.click(link);
fireEvent.click(await screen.findByText("Leave"));
@@ -152,7 +181,7 @@ describe("ImportResumeDialog — Set up a provider", () => {
it("leaves modifier clicks to the browser so the link can open in a new tab", async () => {
renderDialog();
const link = await selectPdfFile();
const link = await selectWordFile();
const event = new MouseEvent("click", { bubbles: true, cancelable: true, metaKey: true });
fireEvent(link, event);
@@ -164,7 +193,7 @@ describe("ImportResumeDialog — Set up a provider", () => {
it("leaves middle clicks to the browser too", async () => {
renderDialog();
const link = await selectPdfFile();
const link = await selectWordFile();
const event = new MouseEvent("click", { bubbles: true, cancelable: true, button: 1 });
fireEvent(link, event);
+42 -15
View File
@@ -141,10 +141,15 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
setIsImporting(true);
// A PDF parsed in the browser never touches a provider, so promising one would be a lie.
const isLocalPdf = value.type === "pdf" && !hasUsableProvider;
const toastId = toast.add({
type: "loading",
title: t`Importing your resume...`,
description: t`This may take a few minutes, depending on the response of the AI provider. Please do not close the window or refresh the page.`,
description: isLocalPdf
? t`This may take a moment. Please do not close the window or refresh the page.`
: t`This may take a few minutes, depending on the response of the AI provider. Please do not close the window or refresh the page.`,
});
try {
@@ -160,14 +165,31 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
if (value.type === "pdf") {
if (isLoadingAiProviders) throw new Error(t`Loading AI providers. Please try again in a moment.`);
if (!hasUsableProvider)
throw new Error(t`This feature requires a connected AI provider. Please set one up in the settings.`);
const base64 = await fileToBase64(value.file);
if (hasUsableProvider) {
const base64 = await fileToBase64(value.file);
data = await client.ai.parsePdf({
file: { name: value.file.name, data: base64 },
});
data = await client.ai.parsePdf({
file: { name: value.file.name, data: base64 },
});
} else {
const [{ extractPdfLines }, { parseResumeText }] = await Promise.all([
import("@/features/resume/import/pdf-text"),
import("@reactive-resume/import/plain-text"),
]);
const lines = await extractPdfLines(value.file);
if (lines.length === 0) {
throw new Error(
t({
comment: "Error shown when a PDF has no extractable text layer during import",
message: "This PDF has no readable text. It is likely a scan, so there is nothing to import.",
}),
);
}
data = parseResumeText(lines.join("\n"));
}
}
if (value.type === "docx") {
@@ -236,7 +258,8 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
const type = useStore(form.store, (s) => s.values.type);
const file = useStore(form.store, (s) => s.values.file);
const aiRequired = type === "pdf" || type === "docx";
const aiRequired = type === "docx";
const pdfWithoutAi = type === "pdf" && !isLoadingAiProviders && !hasUsableProvider;
const onSelectFile = () => {
if (!inputRef.current) return;
@@ -372,12 +395,7 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
{
value: "pdf",
textValue: t({ comment: "File format label in import source selector", message: "PDF" }),
label: (
<div className="flex items-center gap-x-2">
{t({ comment: "File format label in import source selector", message: "PDF" })}{" "}
<Badge>{t`AI`}</Badge>
</div>
),
label: t({ comment: "File format label in import source selector", message: "PDF" }),
},
{
value: "docx",
@@ -413,7 +431,7 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
{aiRequired && !isLoadingAiProviders && !hasUsableProvider && (
<div className="flex flex-col gap-3 rounded-md border border-dashed p-3 text-sm lg:flex-row lg:items-center lg:justify-between">
<span className="text-muted-foreground">
<Trans>Importing from PDF or Word requires a connected AI provider.</Trans>
<Trans>Importing from Word requires a connected AI provider.</Trans>
</span>
<Button
size="sm"
@@ -428,6 +446,15 @@ export function ImportResumeDialog(_: DialogProps<"resume.import">) {
</div>
)}
{pdfWithoutAi && (
<div className="rounded-md border border-dashed p-3 text-muted-foreground text-sm">
<Trans>
No AI provider is connected, so we will read the text out of the PDF here in your browser and fill in what
we can recognize. Expect to tidy up the result.
</Trans>
</div>
)}
<DialogFooter>
<Button type="submit" disabled={!type || !file || isImporting || (aiRequired && !hasUsableProvider)}>
{isImporting ? <Spinner /> : null}
@@ -20,6 +20,8 @@ export type ExtractProgress = HarvestProgress | { phase: "loading"; page: 0; pag
export type ExtractOptions = {
onProgress?: (progress: ExtractProgress) => void;
signal?: AbortSignal;
/** Set to 0 to skip the operator pass; a caller that only wants the text layer has no use for it. */
operatorBudgetMs?: number;
};
export class PdfPasswordRequiredError extends Error {
@@ -81,6 +83,7 @@ export async function extractPdf(file: File, options: ExtractOptions = {}): Prom
return await harvestPdfDocument(document, {
file: { name: file.name, sizeBytes: file.size, magicBytesOk },
maxPages: HARVEST_DEFAULTS.maxPages,
...(options.operatorBudgetMs === undefined ? {} : { operatorBudgetMs: options.operatorBudgetMs }),
...(options.onProgress ? { onProgress: options.onProgress } : {}),
...(options.signal ? { signal: options.signal } : {}),
});
@@ -0,0 +1,117 @@
import type { RawExtraction } from "@reactive-resume/resume/ats-pdf";
import { describe, expect, it } from "vitest";
import { buildExtractedDocument } from "@reactive-resume/resume/ats-pdf";
import { documentToLines } from "./pdf-text";
const PAGE = { width: 595, height: 842 };
const FONT_SIZE = 10;
const CHAR_WIDTH = 5;
type FixtureItem = { text: string; x: number; y: number; width?: number };
/** `y` is measured from the top of the page; the builder flips it into the PDF's own space. */
function rawExtraction(pages: FixtureItem[][]): RawExtraction {
return {
version: 1,
file: { name: "resume.pdf", sizeBytes: 1024, magicBytesOk: true },
pageCount: pages.length,
truncated: false,
pages: pages.map((items, index) => ({
pageNumber: index + 1,
width: PAGE.width,
height: PAGE.height,
rotation: 0,
items: items.map((item) => ({
str: item.text,
transform: [FONT_SIZE, 0, 0, FONT_SIZE, item.x, PAGE.height - item.y - FONT_SIZE] as const,
width: item.width ?? item.text.length * CHAR_WIDTH,
height: FONT_SIZE,
fontRef: "f1",
hasEol: false,
})),
operators: null,
})),
fonts: [],
links: [],
metadata: {
producer: null,
creator: null,
title: null,
author: null,
language: null,
pdfVersion: null,
isEncrypted: false,
isXfa: false,
isCollection: false,
isTagged: false,
hasAcroForm: false,
},
operatorsAvailable: false,
};
}
const linesOf = (...pages: FixtureItem[][]) => documentToLines(buildExtractedDocument(rawExtraction(pages)));
describe("documentToLines", () => {
it("groups items on the same baseline into one line", () => {
expect(
linesOf([
{ text: "Ada", x: 40, y: 60 },
{ text: "Lovelace", x: 58, y: 60 },
]),
).toEqual(["Ada Lovelace"]);
});
it("orders lines from the top of the page down", () => {
expect(
linesOf([
{ text: "Second", x: 40, y: 80 },
{ text: "First", x: 40, y: 60 },
]),
).toEqual(["First", "Second"]);
});
it("turns a wide field gap into a double space the parser can split on", () => {
expect(
linesOf([
{ text: "Acme", x: 40, y: 60 },
{ text: "Engineer", x: 120, y: 60 },
]),
).toEqual(["Acme Engineer"]);
});
it("reads every page in order", () => {
expect(linesOf([{ text: "Page one", x: 40, y: 60 }], [{ text: "Page two", x: 40, y: 60 }])).toEqual([
"Page one",
"Page two",
]);
});
it("drops blank items and blank lines", () => {
expect(
linesOf([
{ text: " ", x: 40, y: 60 },
{ text: "Ada", x: 40, y: 80 },
]),
).toEqual(["Ada"]);
});
it("reads each column of a sidebar layout in full instead of interleaving them", () => {
const sidebar = ["SKILLS", "TypeScript", "Rust", "PostgreSQL", "Docker", "Figma"];
const main = [
"EXPERIENCE",
"Acme Corp",
"Senior Engineer",
"Jan 2020 - Present",
"Led the rewrite",
"Grew the team",
];
expect(
linesOf([
...sidebar.map((text, index) => ({ text, x: 40, y: 60 + index * 20 })),
...main.map((text, index) => ({ text, x: 220, y: 70 + index * 20 })),
]),
).toEqual([...sidebar, ...main]);
});
});
@@ -0,0 +1,82 @@
import type { ExtractedDocument } from "@reactive-resume/resume/ats-pdf";
import { buildExtractedDocument } from "@reactive-resume/resume/ats-pdf";
import { extractPdf } from "@/features/ats-checker/extract-client";
/**
* Turns a PDF into the plain lines `parseResumeText` reads.
*
* The extraction itself is the ATS checker's — same worker configuration, same size and password
* guards, same font-relative line clustering and column-gutter detection. Only the joining differs:
* the ATS text is prose, while the importer splits header rows on runs of two or more spaces, so a
* tabbed "Company Role Location" has to survive as three fields rather than one phrase.
*/
/** Horizontal gap, relative to font size, above which two spans are separate header fields. */
const FIELD_GAP_RATIO = 1.5;
/** Narrower gaps are still a word break; matches the ratio the ATS extractor joins prose with. */
const SPACE_GAP_RATIO = 0.25;
/** The evidence the ATS layout rules require before calling a page two-column. */
const GUTTER_COVERAGE = 0.85;
const GUTTER_SPLIT_RATIO = 0.25;
type PageGeometry = ExtractedDocument["pages"][number];
type TextSpan = PageGeometry["lines"][number]["spans"][number];
function joinSpans(spans: readonly TextSpan[]): string {
let text = "";
let previous: TextSpan | null = null;
for (const span of spans) {
if (previous) {
const gap = span.x - (previous.x + previous.width);
const fontSize = Math.max(previous.fontSize, span.fontSize, 1);
if (gap > FIELD_GAP_RATIO * fontSize) text += " ";
else if (gap > SPACE_GAP_RATIO * fontSize && !/\s$/.test(text) && !/^\s/.test(span.text)) text += " ";
}
text += span.text;
previous = span;
}
return text.trim();
}
/** Mid-point of a convincing gutter, or null on a page that reads as one column. */
function columnSplitX(page: PageGeometry): number | null {
const gutter = page.gutter;
if (!gutter || gutter.coverage < GUTTER_COVERAGE || gutter.splitRatio < GUTTER_SPLIT_RATIO) return null;
return gutter.x + gutter.width / 2;
}
/**
* Lines grouped by baseline alone interleave a sidebar with the body, which drops a whole work
* history into the skills section. On a two-column page each column is read out in full instead.
*/
function pageToLines(page: PageGeometry): string[] {
const splitX = columnSplitX(page);
if (splitX === null) return page.lines.map((line) => joinSpans(line.spans));
const left: string[] = [];
const right: string[] = [];
for (const line of page.lines) {
const leftSpans = line.spans.filter((span) => span.x + span.width / 2 < splitX);
if (leftSpans.length > 0) left.push(joinSpans(leftSpans));
const rightSpans = line.spans.filter((span) => span.x + span.width / 2 >= splitX);
if (rightSpans.length > 0) right.push(joinSpans(rightSpans));
}
return [...left, ...right];
}
export function documentToLines(doc: ExtractedDocument): string[] {
return doc.pages.flatMap(pageToLines).filter((line) => line.length > 0);
}
export async function extractPdfLines(file: File): Promise<string[]> {
return documentToLines(buildExtractedDocument(await extractPdf(file, { operatorBudgetMs: 0 })));
}