Benchmark runs: remove ask_inputs tool, raise child cap to 3h
Sonnet through the Beaver arm paused the banking run on ask_inputs - there is no user in a benchmark run to answer, and the reference harness has no ask-user affordance either (parity). Gate mirrors MIKE_DISABLE_RESEARCH_TOOLS. The 30-min child cap would also have killed any claude-p whole-deliverable run mid-flight (LAB's own sonnet row took 48 min). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018bhP89cYSdEEjHNPnxacns
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 3b712e31 |
| Stats | 2 files changed , +16 , -2 |
| Part of | Chat backend features and prompt/context-budget management |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-9b74c8b3.md
from inside the repo you want the change in.