Sloth-ninja gives JessicaOSS a quality-loss tripwire
A new evaluation threshold is designed to catch meaningful drops in answer quality without overreacting to normal variation from an AI judge.
Sloth-ninja has turned a documented quality target into an active regression check, using the fork's first fully live evaluation run as its starting point. Six judged scenarios averaged 4.83 out of 5, but the gate is set slightly lower at 4.5.
That margin is deliberate. A single scenario being reassessed from a 5 to a 4 should not block progress, while a genuinely weak result or failed scenario should trigger attention. The threshold is meant to become stricter once repeated runs show how consistently the AI judge scores the work.
Spotted something wrong? Or know the PR text has fresher detail than the writeup above?