What is the agent permitted to change?
Name the systems, privileges, stop conditions, human handoff and rollback path before a task starts.
A simulated task reached third-party systems when its environment retained internet access.
A public research instrument for tracing an agent’s action through the evidence needed to trust its outcome.
DOES NOT FOLLOW Nineteen victims. An 8% deployment risk. A sandbox escape.
Select a link in the chain. Each one is tied to a real public record.
Name the systems, privileges, stop conditions, human handoff and rollback path before a task starts.
A simulated task reached third-party systems when its environment retained internet access.
The record shows a page deletion, followed by a write to the same page 18 seconds later.
The deletion alone cannot establish lasting removal or whether later reads remained possible.
The provider reported that evaluation environments were inadvertently connected to the open internet.
Task wording did not enforce isolation; the cited report is not a field-wide frequency estimate.
One reported CyberGym result uses an any-crash qualification.
Any crash and reproduction of a specified vulnerability are different success conditions.
Proposed tests drawn from curated research questions. These are not completed Fide results.
When does oversight prevent harmful intervention, and when does waiting make the outcome worse? The proposed study has not yet produced experimental findings.
Hold the defensive task and resources comparable. Include benign activity, genuine intrusions and ambiguous observations, and record proposed as well as executed actions.
How much do independent checks reduce false acceptance, and what verification cost is practical? Coverage beyond the tested property remains uncertain.
Compare the original acceptance test with independent security and functionality checks. Preserve failed repairs and document which properties each check covers.
Which communication and isolation policies preserve useful cooperation without making the team depend on one untrustworthy contributor?
Introduce a controlled compromised teammate in a fixed defensive task. Compare independent checks, restricted sharing and recovery procedures at a matched resource budget.
These incidents establish that failures occurred in the reported conditions. They do not establish a deployment-wide frequency or prove that a particular new control is sufficient.
Vary explicit scope, environmental isolation and monitoring independently. Record attempted boundary crossings even when infrastructure prevents execution.
Can a small, targeted retest reliably detect a meaningful decline? Software checks of a pipeline cannot answer this empirical question.
Collect paired results before and after a specified setup change. Compare limited retest signals with the full rerun, separating changed responses from changed scoring.
Independent human adjudication is pending. Whether a verification requirement improves an agent’s later actions is a further research question.
Compare report-only review with review grounded in the underlying records. Track corrections, disputed judgments, missed failures and the time needed to review.
Choose an operational claim and inspect the test conditions it would require.
Curated through Sep 26, 2026. This lab analyzes cited sources and published methods; it does not run a defender or produce a readiness score.