docs(lab)+eval: tax gen-8 42/77 - edits land 5/5 & render, coverage gate forces real re-reads; score-neutral on tax too; new `evaluation.receipt` auto-emits per-run score/cost/edit receipt

↗ view on GitHub · Eli Ziff · 2026-08-07 · 5a51b8b1

tax 19-20-25 (echo-v2 + draft-edit-v8) = 42/77 @634k cacheadj, inside
the g7 fixed-treatment band {38,48}. The 5/5 Edit calls on draft.md all
returned ok=true replacements=1 (vs gen-7's dead chars=29), all 5 edit
bodies probe-verified verbatim in the rendered memo.docx (F-15 initial
probe-miss was the en-dash encoding artifact, not absence). Coverage
gate worked end-to-end: first generate_docx refused naming 9 unread
docs → model read exactly those 9 → 5 Edits → bare re-render via
saved draft.md. Mechanism FIXED AND VERIFIED on a second, harder task;
score verdict score-neutral (consistent with HSR) - the lever still
needs an edit-addressable-miss task (acq-style) to show value.

New evaluation/receipt.py: merges scores.json + metrics.json +
beaver-receipts.json into one receipt; `uv run python -m
evaluation.receipt --run-id <id>` prints it (and writes
receipt.json/txt), `--task <t> [--arm a]` prints the stratum table
with base v5 levers filtered so draft_edit/exposure_echo deltas stand
out. run_eval.py now auto-emits the receipt after every judge (both
K=1 and K>1 majority paths).

Co-Authored-By: Claude <noreply@anthropic.com>
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents a7aa5633
Stats 3 files changed , +307
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-5a51b8b1.md from inside the repo you want the change in.

⬇ Download capture-commit-5a51b8b1.md