fix(docx): disable numeric coercion of Word text runs
From the PR description
Summary
Prevent numeric-looking Word text runs from being altered when Mike parses or rewrites DOCX files.
Why / Motivation
fast-xml-parser converts element text to JavaScript numbers and booleans by
default. Word stores visible document text in w:t, so a run containing
12.10 becomes 12.1.
This affects text returned by read_document and can rewrite an untouched run
when Mike saves an unrelated tracked edit.
Changes
- Disable tag-value coercion in the shared DOCX parser.
- Add regression coverage for extracted text and the rewritten
word/document.xml.
Tradeoffs & risks
Element text in word/document.xml is now retained as strings, matching how
the DOCX helper consumes it. Attribute parsing, tracked-change IDs, entity
handling, and whitespace behavior are unchanged.
How verified
- Confirmed the regression test fails on current
main, where the rewrittenw:tcontains12.1. - Confirmed the patched output retains
12.10through an unrelated tracked edit. npm test --prefix backend- 551 passed, 23 skipped.npm run build --prefix backend.
Checklist
- Ran the relevant build and tests.
- Reviewed the diff and removed unrelated changes.
- Docs and environment examples are not applicable.
- No secrets, API keys, real documents, or
.envfiles committed.
Our analysis
Preserve Word text values during DOCX edits — read the full analysis →
Think the analysis missed something the PR description covers?
Commits in this PR (1)
| SHA | Subject | Author | Date | |
|---|---|---|---|---|
c6b3fb3c | fix(docx): disable numeric coercion of Word text runs | Eli Ziff | 2026-08-12 | ↗ GitHub |
Capture this PR into my fork
Download a Markdown prompt that tells Claude how to port every
commit in this PR into your working tree. Run it via
claude -p < capture-pull-328.md from
inside the repo you want the changes in.