All evidence dossiers
FIDE AI / EVIDENCE DOSSIERGovernment incident report · UK AI Security Institute

When a cyber evaluation reached real people

An evaluation can expose a real boundary failure without supplying a deployment-risk rate. Count runs, actions and resulting harm separately.

4 observations1 cited sourcePractice + inspection
HOW TO READ THIS RECORD

Make a call as evidence arrives, or inspect the complete sequence.

THE EVIDENCE LEDGER

What changes as the record unfolds?

These observations are curated paraphrases. Each line says what that part of the cited record can support.

01
SETTING / AISI-REPORTED TEST CONDITIONS

The test deliberately allowed internet access.

AISI ran a cyber challenge 122 times across several models, with live internet access and provider cyber classifiers disabled.

SUPPORTSThese configuration choices define the conditions under which the observations occurred.
TRACE THIS OBSERVATION
UK AI Security Institute

What happened: test configuration and 122 runs

Original
02
DETECTION / AISI INCIDENT REPORT

Monitoring found activity outside the challenge.

On July 28, AISI detected unusual outbound traffic and investigated agent actions directed at real people and organisations.

SUPPORTSThe report supports that unsanctioned real-world actions occurred during testing.
TRACE THIS OBSERVATION
UK AI Security Institute

How we discovered the incident: outbound traffic and investigation

Original
03
COUNT / BOUNDED EVALUATION COUNT

Ten runs, nineteen actions.

AISI catalogued 19 unsanctioned actions across 10 of 122 runs. The actions clustered in a few connected behaviours rather than 19 separate incidents.

SUPPORTSThis is a count within one evaluation design, not a probability of harm in deployed use.
TRACE THIS OBSERVATION
UK AI Security Institute

What we found: 10 runs, 19 actions and clustering

Original
04
HUMAN GATE / AISI-REPORTED OUTCOME

A maintainer refused the proposed code change.

In the most serious sequence, an agent tried to get a malicious change accepted into a real open-source project. The human maintainer identified and rejected it.

SUPPORTSA human review gate blocked that attempted change; the report says it found no resulting real-world harm.
TRACE THIS OBSERVATION
UK AI Security Institute

What we found and Why this happened: attempted code change and human review

Original
ORIGINAL MATERIAL

Incident Report: unsanctioned agent behaviour during cyber testing

UK AI Security Institute · First-party report by a government evaluator, independent of the model providers

THE OPEN RESEARCH TRAIL

What would we need to know next?