fix(pdf): preserve visual layout when extracting PDF text
From the PR description
Problem
extractPdfText joined every pdfjs text item on a page with a blanket space. For real legal documents that mangles the text:
- wrapped clauses lose their line structure (the whole page becomes one flowing string)
- kerning fragments gain spurious spaces (
constitut+e→constitut e) - indentation and column layout (signature blocks, simple tables) is destroyed
Every downstream consumer - chat document context, citation quote matching, anything built from extracted PDF text - inherits that mangled text.
Fix
Rebuild each page's text from the positioned pdfjs items (transform, width, hasEOL) instead of the raw string list:
- Lines are reconstructed from y-coordinates and
hasEOLmarkers, with a y-jump fallback for producers that don't sethasEOLreliably; lines are ordered top-to-bottom, items left-to-right. - Word spacing is driven by measured x-gaps: kerning fragments join without a space, ordinary word gaps get exactly one, and only large gaps become multiple spaces (capped at 16) so signature blocks and simple tables keep their columns without exaggerating justified text.
- Paragraph breaks: a vertical gap above ~1.7× the line height becomes a blank line.
- Indentation is preserved relative to the page's left margin (capped at 24 spaces).
The page-marker format ([Page N]) and the unreadable-buffer `` fallback are unchanged.
Tests
New documentOps.test.ts with a mocked pdfjs covers: line reconstruction from positions, paragraph breaks, column-gap preservation, page markers/ordering, and the failure fallback. Full backend suite passes (607 tests), tsc --noEmit clean.
Our analysis
Preserve PDF layout in extracted text — read the full analysis →
Think the analysis missed something the PR description covers?
Capture this PR into my fork
Download a Markdown prompt that tells Claude how to port every
commit in this PR into your working tree. Run it via
claude -p < capture-pull-336.md from
inside the repo you want the changes in.