Void pilot-a-01 score for arm comparison in LAB protocol
User ruling: the mid-run docker-exec/python-docx failure is our Docker-on-Windows infra fault, so pilot-a-01's 0/95 is not a fair Arm A number. Record it as an incident only; a clean rerun (pilot-a-01r) supplies the comparable number if run. Pilot paused after task 03. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018bhP89cYSdEEjHNPnxacns
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 7b5f9129 |
| Stats | 1 file changed , +5 , -3 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-63c6713f.md
from inside the repo you want the change in.