grounding: Stage 5 pre-registration - scale 4x, cross the checker model
Controls for the composing model being inherently more trustworthy at legal: verification passes can now run on a different model than the composer (finalize gains checkerModel; composition and the repair pass stay on the composer), and the runner gains a --checker-models factor (same | cross | explicit id) recorded per receipt, with control cells running once. Matrix scales to 24 held-constant items: 12 CSLB (4 ordinary case, 4 ordinary legislation, 4 adversarial), 4 CLERC continuations, 8 HousingQA rows (audited 163/0 sufficiency pair plus six unaudited yes/no-balanced rows whose benchmark labels are NOT treated as sufficiency gold). Frozen Hypothesis 5, predictions, and falsification conditions recorded in the log before the run; unit tests 12 passed, tsc clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AgUm5EcRKbT3duVFMbrQD
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 413f7032 |
| Stats | 3 files changed , +118 , -6 |
| Part of | Legal grounding and retrieval research |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-29bc7e6f.md
from inside the repo you want the change in.