Ground-up restart plan: four audits, four failure classes, and the ideas that were never fairly tested

↗ view on GitHub · Eli Ziff · 2026-07-31 · e347b85c

Four parallel audits (assets+perf, case-law discards, retrieval discards, bed
design) found four distinct failure classes, each of which produced confident
verdicts that favoured the incumbent: instrument defects (CRLF, D2 double-count,
D3 fallback blending), statistics borrowed from the wrong n (premise #1 is 8
cells over one item pair; flat_recency retired on p=0.264; H18 falsified on
<=1-cell deltas), detector blindness read as a negative result (zero subsections
across all 69 documents, so R1 never tested clause chunking on the sources where
it failed), and a bookkeeping error that crowned the current composer lane.

Plan phases: P0 stop the bleeding (perDocCap confound blocks every composed run;
ensurePassageIndex rebuilds on transient SQLITE_BUSY; repair the skeleton
detector; withdraw R1). P1 free re-tests on banked data (checker-family crossing
~150 calls decides premise #1; C4 marginal-value study zero calls). P2 build the
case-law bed - gated on the A2AJ full-text import, since the provider sqlite is
metadata_only with 0 of 248,536 documents carrying text. P3 the arms never run,
each shipping the control that makes it falsifiable. P4 close out LegalBench
cheaply and stop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H9ToHYJVDxfeJcJwdzrP2H
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents bd8bb49d
Stats 1 file changed , +267
Part of Legal agent-surface and context-compaction experiments

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-e347b85c.md from inside the repo you want the change in.

⬇ Download capture-commit-e347b85c.md