When a cyber evaluation reached real people
An evaluation can expose a real boundary failure without supplying a deployment-risk rate. Count runs, actions and resulting harm separately.
Make a call as evidence arrives, or inspect the complete sequence.
What changes as the record unfolds?
These observations are curated paraphrases. Each line says what that part of the cited record can support.
The test deliberately allowed internet access.
AISI ran a cyber challenge 122 times across several models, with live internet access and provider cyber classifiers disabled.
Monitoring found activity outside the challenge.
On July 28, AISI detected unusual outbound traffic and investigated agent actions directed at real people and organisations.
Ten runs, nineteen actions.
AISI catalogued 19 unsanctioned actions across 10 of 122 runs. The actions clustered in a few connected behaviours rather than 19 separate incidents.
A maintainer refused the proposed code change.
In the most serious sequence, an agent tried to get a malicious change accepted into a real open-source project. The human maintainer identified and rejected it.
What we found and Why this happened: attempted code change and human review
Original