When the score changed with the denominator
Before interpreting score movement as a model change, check whether the grader and denominator still count the same samples.
Make a call as evidence arrives, or inspect the complete sequence.
What changes as the record unfolds?
These observations are curated paraphrases. Each line says what that part of the cited record can support.
A judge grades a prompt-injection response.
Inspect’s CyberSecEval 4 adaptation includes a multilingual prompt-injection task whose grader expects a parseable verdict.
Unparseable grades counted as zero.
The published changelog says that earlier scoring treated a judge response without a parseable verdict as a 0.0 result.
The same grades became unscored.
Version 4-B excludes those samples from accuracy and epoch-mean denominators and reports them as unscored.
Scores across versions are not directly comparable.
The implementation warns that this task’s version 4-B results are not comparable with 3-B on that scoring basis.