Correct the cross-reference verdict: unproven on this bed, not refuted
The earlier "retrieval-time edge following is DEAD" conclusion was an overreach and is withdrawn. Four reasons the measurement cannot carry it: 1. The metric structurally cannot reward the mechanism. LegalBench gold is where the answer text is, not what must be read to understand it, so resolving "as defined in Section 1.1" to fetch a deal-term definition - correct legal behaviour - scores as precision loss by construction. 2. Small n: 15/168 contractnli and 14/153 cuad gold fragments had an edge at all. The "cuad loses 6x" headline is fourteen fragments. 3. Half instrument strength: 24/69 documents refused by the integrity gate, including 8/17 maud, and the resolver's ceiling (the skeleton section inventory) was only repaired mid-run. 4. It measured a proxy - pool expansion scored by character overlap - not the registered proposal, which is structure as orientation for a composer. Edge-following is therefore UNPROVEN here: it earns no arm yet, and no positive claim either. The similarity-edge retirement stands, since those lost to a deterministic control and do not depend on the gold-span objection. The resolver stands on its own as a working legal primitive. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H9ToHYJVDxfeJcJwdzrP2H
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | a738ae35 |
| Stats | 1 file changed , +42 , -18 |
| Part of | Legal agent-surface and context-compaction experiments |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-0d5e4568.md
from inside the repo you want the change in.