Void pilot-a-01 score for arm comparison in LAB protocol

↗ view on GitHub · Eli Ziff · 2026-07-28 · 63c6713f

User ruling: the mid-run docker-exec/python-docx failure is our
Docker-on-Windows infra fault, so pilot-a-01's 0/95 is not a fair
Arm A number. Record it as an incident only; a clean rerun
(pilot-a-01r) supplies the comparable number if run. Pilot paused
after task 03.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018bhP89cYSdEEjHNPnxacns
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents 7b5f9129
Stats 1 file changed , +5 , -3
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-63c6713f.md from inside the repo you want the change in.

⬇ Download capture-commit-63c6713f.md