Stage 17 verdict: crowned stack holds on full LegalBench-RAG mini (776 cells)
651 answered (83.9%), P=0.505 answered-only grounded char precision vs 0.087 baseline at 1/7th the quoted chars; 1,358/1,358 quoted spans verbatim-audited, zero false passes; doc-miss audit resolves 6/7 to a byte-identical CUAD duplicate file, true rate 0.15%. Runner: fix usage-undefined crash in the finalizer merge; add --rerank-preview so a resume never mixes rerank configs in one receipt file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AgUm5EcRKbT3duVFMbrQD
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 1470f4f4 |
| Stats | 2 files changed , +74 , -7 |
| Part of | Legal grounding and retrieval research |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-25ef6ec5.md
from inside the repo you want the change in.