fix: extract bolded verdicts in legalbench label extraction
Chat models bold their verdict inside verbose analyses; the whole-text fallback saw both labels and refused. Bold spans now count when they map to exactly one label. Rescoring the saved codex sweep flips citation_prediction_classification from 31.7 (16 unparsed) to 88.1 (2). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016yXxuitxX5h4wm1v88hnga
| Repository | eliziff/Beaver |
|---|---|
| Author | Eli Ziff <eliasziff@gmail.com> |
| Authored | |
| Parents | 4c9929dd |
| Stats | 2 files changed , +22 |
| Part of | Evaluation harness: Beaver-CAN and LegalBench-RAG adapters |
Capture this commit into my fork
Download a Markdown prompt that tells Claude how to port this
exact commit into your working tree. Run it via
claude -p < capture-commit-361191e1.md
from inside the repo you want the change in.