FIDE AI / EVIDENCE DOSSIERSource comparison · 2 cited sources
Two success metrics, two different questions
Before comparing scores, compare what counts as success. An impressive result can still answer a different question.
4 observations2 cited sourcesInspection
THE EVIDENCE LEDGER
What changes as the record unfolds?
These observations are curated paraphrases. Each line says what that part of the cited record can support.
01
TASK / BENCHMARK DEFINITION
Reproduce a specified vulnerability.
CyberGym evaluates vulnerability reproduction from a description and an unpatched codebase.
SUPPORTSRead the benchmark task and validation criteria.
TRACE THIS OBSERVATION
02
CLAIM / VENDOR-REPORTED METRIC
A product announcement reports an any-crash result.
Microsoft explicitly qualifies its headline CyberGym score as an any-crash score.
SUPPORTSThe metric qualification is part of the claim, not a footnote to discard.
03
COMPARISON / EDITORIAL INTERPRETATION
A crash may not be the target vulnerability.
Triggering some crash and reproducing the specified vulnerability have different success conditions.
SUPPORTSDo not rank unlike measurements as if they were interchangeable.
TRACE THIS OBSERVATION
04
NEXT CHECK / SUGGESTED REVIEW PRACTICE
Match definitions before comparing.
Record benchmark version, validation rule, agent configuration and resource budget.
SUPPORTSIf these details are unavailable, mark the comparison as incomplete.
TRACE THIS OBSERVATION