Vendor LAB dev+validation tiers and the generalization corpus in-repo
benchmarks/harvey-labs carries the 210 visible tasks, the grading harness with its 13 integration patches (provenance and upstream HEAD in PROVENANCE.md), and our run results; the 997 sealed tasks exist only upstream. benchmarks/legal-generalization-corpus is the held-out gold corpus the scanners validate against. Local venv and caches stay untracked. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AgUm5EcRKbT3duVFMbrQD
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | f5297816 |
| Stats | 18527 files changed , +1402582 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-d43fec5f.md
from inside the repo you want the change in.