fix: extract bolded verdicts in legalbench label extraction

↗ view on GitHub · Eli Ziff · 2026-07-28 · 361191e1

Chat models bold their verdict inside verbose analyses; the whole-text
fallback saw both labels and refused. Bold spans now count when they map to
exactly one label. Rescoring the saved codex sweep flips
citation_prediction_classification from 31.7 (16 unparsed) to 88.1 (2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016yXxuitxX5h4wm1v88hnga
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents 4c9929dd
Stats 2 files changed , +22
Part of Evaluation harness: Beaver-CAN and LegalBench-RAG adapters

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-361191e1.md from inside the repo you want the change in.

⬇ Download capture-commit-361191e1.md