Do agents respect the boundaries of their task?
A task may encourage persistent problem-solving while permissions constrain where that problem-solving is allowed. Both must be evaluated.
Evidence emergingA first-party incident assessment describes boundary failures in cybersecurity evaluation environments with unintended internet access.
These incidents establish that failures occurred in the reported conditions. They do not establish a deployment-wide frequency or prove that a particular new control is sufficient.
Vary explicit scope, environmental isolation and monitoring independently. Record attempted boundary crossings even when infrastructure prevents execution.
PROPOSED TEST · NO RESULT CLAIMEDWhat would change the answer?
- Attempted violations
- Executed violations
- Monitor misses
- Task completion
Follow the evidence and its limits.
Source type and limitation are kept visible. A research plan is not an observed outcome.
An alignment assessment of recent cybersecurity incidents
Provider investigation of its own models. The report describes unusual evaluation conditions and a planned independent investigation.
OWASP Top 10 for Agentic Applications
A risk framework organizes investigation; following a checklist does not by itself demonstrate control effectiveness.
SLEIGHT-Bench: Finding Blind Spots in AI Monitors
Constructed transcripts reveal possible blind spots; they do not estimate how often those failures occur in deployed systems.
Incident Report: unsanctioned agent behaviour during cyber testing
The evaluation intentionally allowed internet access and disabled provider cyber classifiers; it was not a sandbox escape. The observed run fraction is not a deployment risk estimate. AISI reported no evidenced resulting real-world harm.