feat: evaluation harness - traces, Beaver-CAN, validators, four-arm runner

↗ view on GitHub · Eli Ziff · 2026-07-27 · b0a22722

Eval plan Issues 1-4: strict run-trace contract with git/config/token/
cost provenance; Beaver-CAN task+gold schemas with a three-task dev
slice whose pinpoints are validated by the production source compiler;
deterministic validator primitives (quotations, packet-only sources,
seeded-identifier leakage, headings, provenance IDs, filenames, DOCX
revision/comment structure); and a four-arm runner (bare_model,
oracle_sources, beaver_baseline, beaver_candidate) producing per-arm
validated traces and a comparison report. Unscored criteria are
explicit nulls, never faked. streamResponsesApi now reports provider
token usage so traces carry real tokens and cost. Verified with a live
two-arm gpt-5-mini run on CAN-RETRIEVAL-001.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yXxuitxX5h4wm1v88hnga
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents 2512207f
Stats 31 files changed , +3448 , -1
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-b0a22722.md from inside the repo you want the change in.

⬇ Download capture-commit-b0a22722.md