Fix stress metric: compare base-label multisets, not @n-suffixed labels

↗ view on GitHub · Eli Ziff · 2026-07-28 · 53831b95

Occurrence suffixes disambiguate restarts within one parse but cascade
under damage, making cross-version label recall non-monotonic (the
CBCA caveat from the original harness run). Comparing base-label
multisets restores a meaningful curve - ON ESA 2000: 95.9/90.7/85.2/
72.6/55.2% across 0.2-5% corruption, misplacement 97.1% salvage with
violations rising as the disorder signal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018bhP89cYSdEEjHNPnxacns
Repository eliziff/Beaver
Author Eli Ziff <eliasziff@gmail.com>
Authored
Parents f001b213
Stats 1 file changed , +23 , -8
Part of Deterministic legal-text tools ("SLA audit organs")

Capture this commit into my fork

Download a Markdown prompt that tells Claude how to port this exact commit into your working tree. Run it via claude -p < capture-commit-53831b95.md from inside the repo you want the change in.

⬇ Download capture-commit-53831b95.md