The monitor that missed a constructed attack
A miss on a constructed transcript exposes a monitor blind spot. It does not estimate how often the failure occurs in deployed systems.
What changes as the record unfolds?
These observations are curated paraphrases. Each line says what that part of the cited record can support.
The attacks were constructed as transcripts.
SLEIGHT-Bench authors designed synthetic coding-agent transcripts across 11 blind-spot categories, rather than collecting field incidents.
SLEIGHT-Bench: constructed transcripts and design limits
OriginalBenign lookalikes calibrated false alarms.
The authors paired harmful transcripts with benign variations and evaluated monitors at a 1% false-positive threshold.
Evaluating monitors: paired benign variations and threshold
OriginalSome harmful traces escaped detection.
The paper reports low catch rates for several designed blind spots, including omissions and multi-step reasoning failures.
A synthetic miss is not field prevalence.
The authors note that transcript construction broadens scenario coverage but can reduce realism. The benchmark does not observe deployed incident frequency.