EVIDENCE TRIALS / 4 PUBLIC RECORDS

Read the record. Test the claim.

Practice the distinctions that matter in autonomous cyber defense. Each dossier is based on a linked public account; every judgment has a stated limit.

Choose a record. Each has two judgments. Your progress stays in this browser.

YOUR REASONING RECORD
0 / 8 judgments
FIDE AI / EVIDENCE DOSSIERGovernment incident report · UK AI Security Institute

When a cyber evaluation reached real people

An evaluation can expose a real boundary failure without supplying a deployment-risk rate. Count runs, actions and resulting harm separately.

4 observations1 cited sourcePractice + inspection
HOW TO READ THIS RECORD

Make a call as evidence arrives, or inspect the complete sequence.

THE RECORD / IN SEQUENCE
E01

AISI ran one cyber challenge 122 times across several models.

UK AI Security Institute report
E02

It found unsanctioned live-internet actions in 10 runs and catalogued 19 actions.

UK AI Security Institute report
PROVENANCEUK AI Security Institute

First-party report by a government evaluator, independent of the model providers

Source note
CLAIM UNDER INSPECTIONDISPATCH 01 / 02

Which conclusion is supported by this count?

BOUNDARY OF THE SOURCE

The evaluation intentionally allowed internet access and disabled provider cyber classifiers; it was not a sandbox escape. The observed run fraction is not a deployment risk estimate. AISI reported no evidenced resulting real-world harm.

ORIGINAL MATERIAL

Incident Report: unsanctioned agent behaviour during cyber testing

UK AI Security Institute · First-party report by a government evaluator, independent of the model providers

THE OPEN RESEARCH TRAIL

What would we need to know next?