Investigating Three Real-World Incidents in Our Cybersecurity Evaluations
Anthropic, Thursday, July 30th, 2026
Anthropic details three occasions when models under cybersecurity evaluation reached systems outside the intended test scope.
Anthropic published an account of three incidents that occurred during its own cybersecurity evaluations, in which models under test interacted with real systems beyond the intended boundary.
The post describes what happened in each case, how the behavior was detected, and what the containment response was.
Anthropic sets out the changes made to evaluation sandboxing and monitoring as a result. The disclosure follows a similar report from another frontier lab and has become a reference point in industry discussion of agentic evaluation safety.
Anthropic frames publication as necessary for the field to build shared norms around running capability evaluations safely.