Cases as evidence,
not anecdotes.
Each dossier follows the observations, the claim they can support, and the questions still outside the record. Practice is available where a bounded judgment has been authored.
The cleanup that needed another check
Check the outcome relevant to the decision. Removing an artifact and establishing durable recovery require different evidence.
When a simulated task reached real systems
A task’s framing cannot substitute for enforced isolation. Keep behavioral failures distinct from the environment conditions that made them possible.
Two success metrics, two different questions
Before comparing scores, compare what counts as success. An impressive result can still answer a different question.
When a cyber evaluation reached real people
An evaluation can expose a real boundary failure without supplying a deployment-risk rate. Count runs, actions and resulting harm separately.
The monitor that missed a constructed attack
A miss on a constructed transcript exposes a monitor blind spot. It does not estimate how often the failure occurs in deployed systems.
When the score changed with the denominator
Before interpreting score movement as a model change, check whether the grader and denominator still count the same samples.
These are curated paraphrases of cited sources, not full transcripts or claims of independent reproduction.