All evidence dossiers
FIDE AI / EVIDENCE DOSSIERBenchmark finding · Anthropic Fellows, Redwood & Anthropic

The monitor that missed a constructed attack

A miss on a constructed transcript exposes a monitor blind spot. It does not estimate how often the failure occurs in deployed systems.

4 observations1 cited sourceInspection
THE EVIDENCE LEDGER

What changes as the record unfolds?

These observations are curated paraphrases. Each line says what that part of the cited record can support.

01
SETTING / PUBLISHED BENCHMARK DESIGN

The attacks were constructed as transcripts.

SLEIGHT-Bench authors designed synthetic coding-agent transcripts across 11 blind-spot categories, rather than collecting field incidents.

SUPPORTSThis is a stress test of possible monitor failure modes.
TRACE THIS OBSERVATION
Anthropic Fellows, Redwood & Anthropic

SLEIGHT-Bench: constructed transcripts and design limits

Original
02
COMPARISON / PUBLISHED MEASUREMENT METHOD

Benign lookalikes calibrated false alarms.

The authors paired harmful transcripts with benign variations and evaluated monitors at a 1% false-positive threshold.

SUPPORTSDetection has to be understood alongside false-positive cost.
TRACE THIS OBSERVATION
Anthropic Fellows, Redwood & Anthropic

Evaluating monitors: paired benign variations and threshold

Original
03
FINDING / REPORTED BENCHMARK RESULT

Some harmful traces escaped detection.

The paper reports low catch rates for several designed blind spots, including omissions and multi-step reasoning failures.

SUPPORTSThe tested monitors missed particular constructed attacks under the reported prompts and threshold.
TRACE THIS OBSERVATION
04
LIMIT / TRANSFER BOUNDARY

A synthetic miss is not field prevalence.

The authors note that transcript construction broadens scenario coverage but can reduce realism. The benchmark does not observe deployed incident frequency.

SUPPORTSUse the result to design further tests, not to claim an operational failure rate.
TRACE THIS OBSERVATION
Anthropic Fellows, Redwood & Anthropic

SLEIGHT-Bench design limits and Conclusion

Original
THE OPEN RESEARCH TRAIL

What would we need to know next?