Vendor LAB dev+validation tiers and the generalization corpus in-repo

↗ view on GitHub · Eli Ziff · 2026-07-29 · d43fec5f

benchmarks/harvey-labs carries the 210 visible tasks, the grading
harness with its 13 integration patches (provenance and upstream HEAD in
PROVENANCE.md), and our run results; the 997 sealed tasks exist only
upstream. benchmarks/legal-generalization-corpus is the held-out gold
corpus the scanners validate against. Local venv and caches stay
untracked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AgUm5EcRKbT3duVFMbrQD
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents f5297816
Stats 18527 files changed , +1402582
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-d43fec5f.md from inside the repo you want the change in.

⬇ Download capture-commit-d43fec5f.md