feat(lab): lab-compare - paired per-criterion arm comparison

↗ view on GitHub · Eli Ziff · 2026-08-05 · 1b4ecebb

Never compare score totals again: per-criterion majority verdicts across
replicates, McNemar exact binomial per task + pooled (with the
pseudo-replication caveat printed), per-run metrics rows, price-weighted
cost C = uncached + 0.1*cache_read + r*output at r in {1,2,4,6}, and the
deliverable-length/pass-rate correlation (the verbosity confound). Refuses
cross-judge comparisons. Prefers scores.majority.json over scores.json.
Verified against the existing HSR deepseek runs (reproduces the audit's
b=1/c=4, p=0.375 e2e-vs-index contrast).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011pfUVhNFTRvhYGXBwoKNj6
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents d00f2f1e
Stats 1 file changed , +239
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-1b4ecebb.md from inside the repo you want the change in.

⬇ Download capture-commit-1b4ecebb.md