All evidence dossiers
FIDE AI / EVIDENCE DOSSIERSource comparison · 2 cited sources

Two success metrics, two different questions

Before comparing scores, compare what counts as success. An impressive result can still answer a different question.

4 observations2 cited sourcesInspection
THE EVIDENCE LEDGER

What changes as the record unfolds?

These observations are curated paraphrases. Each line says what that part of the cited record can support.

01
TASK / BENCHMARK DEFINITION

Reproduce a specified vulnerability.

CyberGym evaluates vulnerability reproduction from a description and an unpatched codebase.

SUPPORTSRead the benchmark task and validation criteria.
TRACE THIS OBSERVATION
UC Berkeley researchers

Overview of CyberGym: vulnerability reproduction task

Original
02
CLAIM / VENDOR-REPORTED METRIC

A product announcement reports an any-crash result.

Microsoft explicitly qualifies its headline CyberGym score as an any-crash score.

SUPPORTSThe metric qualification is part of the claim, not a footnote to discard.
TRACE THIS OBSERVATION
Microsoft

A multi-model architecture built for security; any-crash note

Original
03
COMPARISON / EDITORIAL INTERPRETATION

A crash may not be the target vulnerability.

Triggering some crash and reproducing the specified vulnerability have different success conditions.

SUPPORTSDo not rank unlike measurements as if they were interchangeable.
TRACE THIS OBSERVATION
UC Berkeley researchers

Overview of CyberGym: target vulnerability success check

Original
Microsoft

Any-crash score qualification

Original
04
NEXT CHECK / SUGGESTED REVIEW PRACTICE

Match definitions before comparing.

Record benchmark version, validation rule, agent configuration and resource budget.

SUPPORTSIf these details are unavailable, mark the comparison as incomplete.
TRACE THIS OBSERVATION
UC Berkeley researchers

Benchmark task and validation rule

Original
Microsoft

Reported any-crash score qualification

Original
THE OPEN RESEARCH TRAIL

What would we need to know next?