Structure-graph round: the kill criterion fired, and it killed the retrieval arm
Both deterministic measurements ran before any model-side arm, as registered. The resolver works (9,234 references detected, 2,428 resolved, 3.3% unresolved, 24/69 documents refused by the integrity gate) and its ceiling is provably the skeleton's section inventory - maud docs with 97-111 sections resolve 96-97%, docs with 13-33 resolve 24-40%. But the edges do not point at gold. Against a contiguous same-budget control, graph precision loses on contractnli (4.06 vs 5.06) and loses 6x on cuad (1.15 vs 6.99); maud is the only win at 1.6x, and reaches just 13.7% of its fragmented gold. Same shape as the C1 coverage arms. So retrieval-time edge following - argued in session as the cheapest and strongest version of the idea - is DEAD on this bed and gets no arm. The weak layers are retired on evidence: defined-term 0.12-1.29% and lexical 0.81-4.57%, below the contiguous control everywhere. The surviving hypothesis is narrower: the graph as a map a composer corrects, i.e. orientation rather than evidence, to be registered and gated on its own terms. Zero model calls spent reaching this. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H9ToHYJVDxfeJcJwdzrP2H
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 2d903c82 |
| Stats | 1 file changed , +55 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-a738ae35.md
from inside the repo you want the change in.