docs: adopt the minimal evaluation and context plan

↗ view on GitHub · Eli Ziff · 2026-07-27 · aaa86ecd

Measurement-first scoping of the external plan against what the repo
already owns: extend contextManifest into a common eval trace first,
CanLegalRAGBench retrieval-stage second, LongMemEval when a second
context strategy exists. LegalBench-RAG-mini deferred (no chunk-RAG
exists to measure) and the maximal brief's four-arm portfolio rejected
as infrastructure-in-the-dark. OpenAI prompt caching recorded as already
measured (2,304-token stable prefix reuse across the conditional
spreadsheet splice).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QUsMXzBsEJ27DjJVPSfW5o
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents f2e585ba
Stats 1 file changed , +160
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-aaa86ecd.md from inside the repo you want the change in.

⬇ Download capture-commit-aaa86ecd.md