Arm B becomes the whole bet: disclose tools on demand, and let the document say what it affords

↗ view on GitHub · Eli Ziff · 2026-07-31 · 6eec3f04

Eli: "none of the existing stuff is really tested. I am not afraid about
testing multiple things at once. Make B our best guess at combining all of
our ideas." So B stops being a naming change and becomes the full vision,
measured against a frozen A.

PROGRESSIVE DISCLOSURE. Tool-selection accuracy is measured to degrade past
roughly 20-25 tools and this surface carries 44, which is very likely a
larger effect than any wording I have been trimming. B keeps the eleven verbs
a task starts with -- list, outline, read, find, links, lookup, evidence,
apply_text_ops, ask, submit, describe_tools -- and defers 34 behind one
discovery call, grouped by how a task arrives (research, drafting, review,
amendment, authorities, deadlines, workflow) rather than by module.

  resident schema   A 10,130 tokens   B 4,004 tokens   -60.5%

That is the answer to Eli's "shouldn't the new one be SHORTER?" -- it is,
once the lever is the tool count rather than the prose. Trimming descriptions
was worth ~80 tokens; this is worth 6,126.

CAPABILITY ON CONTACT. A schema can only describe addressing in the abstract,
identically for a paginated PDF and a DOCX with no pages. The document knows.
The opening read now reports what THIS document affords -- sections, page
count and which schemes address them, resolved cross-references, contents
outline, or a note saying there is no numbering -- so the model asks for
things that exist. Only on the opening read: it costs a skeleton compile, and
repeating it per windowed read would pay that every turn for information that
has not changed.

The known risk, to be measured rather than assumed: a model does not call
what it cannot see, and the prompt currently PUSHES the deterministic organs.
If B under-calls them, that is a quality regression a token count will not
show, and it disqualifies the arm regardless of the saving.

partitionTools/toolsForDomains are exported because the harness owns the
conversation loop and has to add revealed tools to the next request.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H9ToHYJVDxfeJcJwdzrP2H
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents 538072b0
Stats 1 file changed , +190
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-6eec3f04.md from inside the repo you want the change in.

⬇ Download capture-commit-6eec3f04.md