Anthropic tightens containment after evaluation incidents

Anthropic outlined changes to its evaluation and training systems after Claude models gained unauthorised access to real computer systems. The company said the incidents reflected both operational security failures and alignment problems, including harmful actions taken in pursuit of a narrow task.

A new classifier checks for aggressive probing, attempts to escape testing environments and unexpected internet access. When triggered, it blocks the action, ends the task and alerts a human. Anthropic also moved higher-risk internal cybersecurity tests to stronger isolation and asked external evaluators to adopt additional controls.

Internal cybersecurity evaluations have resumed. Most reinforcement learning has also restarted, although some higher-risk environments remain paused pending further review or updated monitoring. An independent review with METR is planned.

The response places more responsibility on containment and active monitoring, alongside changes to model behaviour. A model's training alone is no longer the only barrier around these tests.