Does a convincing report faithfully represent its evidence?
A polished explanation can contain both useful findings and an unsupported conclusion. Review the claims that determine the decision.
Published analysisFide’s DSEWiki assessment examines whether follow-up reports corrected earlier claims, separately from overall benchmark performance.
Independent human adjudication is pending. Whether a verification requirement improves an agent’s later actions is a further research question.
Compare report-only review with review grounded in the underlying records. Track corrections, disputed judgments, missed failures and the time needed to review.
PROPOSED TEST · NO RESULT CLAIMEDWhat would change the answer?
- Unsupported claims
- Corrections
- Reviewer agreement
- Review cost
Follow the evidence and its limits.
Source type and limitation are kept visible. A research plan is not an observed outcome.
DSEWiki: What the records showed, and AI reports missed
Assistant-coded assessments; independent human adjudication is pending. Findings describe this report collection, not all AI investigators.
Inspect: a framework for language model evaluations
An evaluation framework provides infrastructure; the validity of conclusions still depends on task and scoring design.
SLEIGHT-Bench: Finding Blind Spots in AI Monitors
Constructed transcripts reveal possible blind spots; they do not estimate how often those failures occur in deployed systems.
CyberSecEval 4 in Inspect Evals
Autonomous-uplift and autopatching prototypes are excluded. Reported subset results do not establish deployment readiness.
Incident Report: unsanctioned agent behaviour during cyber testing
The evaluation intentionally allowed internet access and disabled provider cyber classifiers; it was not a sandbox escape. The observed run fraction is not a deployment risk estimate. AISI reported no evidenced resulting real-world harm.