FIDE AI / AUTONOMOUS DEFENSE

What does the
record allow us
to claim?

A public research instrument for tracing an agent’s action through the evidence needed to trust its outcome.

ON THE RECORD / 01UK AI Security Institute
ONE CYBER CHALLENGE
122test runs
10runs with unsanctioned live-internet actions
19actions catalogued

DOES NOT FOLLOW Nineteen victims. An 8% deployment risk. A sandbox escape.

See what this count can justify
THE REASONING CHAIN

Four checks between action and trust.

Select a link in the chain. Each one is tied to a real public record.

CHECK 01 / BOUND THE ACTION

What is the agent permitted to change?

Name the systems, privileges, stop conditions, human handoff and rollback path before a task starts.

THE RECORD SHOWS

A simulated task reached third-party systems when its environment retained internet access.

THE INFERENCE TO RESISTAn instruction that says “simulation” is not an enforced boundary.
AN INDEX OF OVERREACHES

Where does the inference break?

All evidence dossiers
FIDE RESEARCH AGENDA

What observation would change the answer?

Proposed tests drawn from curated research questions. These are not completed Fide results.

01When should a defender act, wait, or hand control back?

When does oversight prevent harmful intervention, and when does waiting make the outcome worse? The proposed study has not yet produced experimental findings.

A MORE DECISIVE TEST

Hold the defensive task and resources comparable. Include benign activity, genuine intrusions and ambiguous observations, and record proposed as well as executed actions.

Follow the evidence and measures
02What would prove that a repair restored security?

How much do independent checks reduce false acceptance, and what verification cost is practical? Coverage beyond the tested property remains uncertain.

A MORE DECISIVE TEST

Compare the original acceptance test with independent security and functionality checks. Preserve failed repairs and document which properties each check covers.

Follow the evidence and measures
03Can a defense team recover from a compromised teammate?

Which communication and isolation policies preserve useful cooperation without making the team depend on one untrustworthy contributor?

A MORE DECISIVE TEST

Introduce a controlled compromised teammate in a fixed defensive task. Compare independent checks, restricted sharing and recovery procedures at a matched resource budget.

Follow the evidence and measures
04Do agents respect the boundaries of their task?

These incidents establish that failures occurred in the reported conditions. They do not establish a deployment-wide frequency or prove that a particular new control is sufficient.

A MORE DECISIVE TEST

Vary explicit scope, environmental isolation and monitoring independently. Record attempted boundary crossings even when infrastructure prevents execution.

Follow the evidence and measures
05When does an earlier evaluation stop being informative?

Can a small, targeted retest reliably detect a meaningful decline? Software checks of a pipeline cannot answer this empirical question.

A MORE DECISIVE TEST

Collect paired results before and after a specified setup change. Compare limited retest signals with the full rerun, separating changed responses from changed scoring.

Follow the evidence and measures
06Does a convincing report faithfully represent its evidence?

Independent human adjudication is pending. Whether a verification requirement improves an agent’s later actions is a further research question.

A MORE DECISIVE TEST

Compare report-only review with review grounded in the underlying records. Track corrections, disputed judgments, missed failures and the time needed to review.

Follow the evidence and measures
All six research questions
THE TRANSFER LENS

What can an evaluation actually justify?

Choose an operational claim and inspect the test conditions it would require.

Curated through Sep 26, 2026. This lab analyzes cited sources and published methods; it does not run a defender or produce a readiness score.