docs: adopt the minimal evaluation and context plan
Measurement-first scoping of the external plan against what the repo already owns: extend contextManifest into a common eval trace first, CanLegalRAGBench retrieval-stage second, LongMemEval when a second context strategy exists. LegalBench-RAG-mini deferred (no chunk-RAG exists to measure) and the maximal brief's four-arm portfolio rejected as infrastructure-in-the-dark. OpenAI prompt caching recorded as already measured (2,304-token stable prefix reuse across the conditional spreadsheet splice). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QUsMXzBsEJ27DjJVPSfW5o
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | f2e585ba |
| Stats | 1 file changed , +160 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-aaa86ecd.md
from inside the repo you want the change in.