feat(lab): lab-compare - paired per-criterion arm comparison
Never compare score totals again: per-criterion majority verdicts across
replicates, McNemar exact binomial per task + pooled (with the
pseudo-replication caveat printed), per-run metrics rows, price-weighted
cost C = uncached + 0.1*cache_read + r*output at r in {1,2,4,6}, and the
deliverable-length/pass-rate correlation (the verbosity confound). Refuses
cross-judge comparisons. Prefers scores.majority.json over scores.json.
Verified against the existing HSR deepseek runs (reproduces the audit's
b=1/c=4, p=0.375 e2e-vs-index contrast).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011pfUVhNFTRvhYGXBwoKNj6
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | d00f2f1e |
| Stats | 1 file changed , +239 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-1b4ecebb.md
from inside the repo you want the change in.