docs(lab)+eval: tax gen-8 42/77 - edits land 5/5 & render, coverage gate forces real re-reads; score-neutral on tax too; new `evaluation.receipt` auto-emits per-run score/cost/edit receipt
tax 19-20-25 (echo-v2 + draft-edit-v8) = 42/77 @634k cacheadj, inside
the g7 fixed-treatment band {38,48}. The 5/5 Edit calls on draft.md all
returned ok=true replacements=1 (vs gen-7's dead chars=29), all 5 edit
bodies probe-verified verbatim in the rendered memo.docx (F-15 initial
probe-miss was the en-dash encoding artifact, not absence). Coverage
gate worked end-to-end: first generate_docx refused naming 9 unread
docs → model read exactly those 9 → 5 Edits → bare re-render via
saved draft.md. Mechanism FIXED AND VERIFIED on a second, harder task;
score verdict score-neutral (consistent with HSR) - the lever still
needs an edit-addressable-miss task (acq-style) to show value.
New evaluation/receipt.py: merges scores.json + metrics.json +
beaver-receipts.json into one receipt; `uv run python -m
evaluation.receipt --run-id <id>` prints it (and writes
receipt.json/txt), `--task <t> [--arm a]` prints the stratum table
with base v5 levers filtered so draft_edit/exposure_echo deltas stand
out. run_eval.py now auto-emits the receipt after every judge (both
K=1 and K>1 majority paths).
Co-Authored-By: Claude <noreply@anthropic.com>
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | a7aa5633 |
| Stats | 3 files changed , +307 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-5a51b8b1.md
from inside the repo you want the change in.