Claude agents improve models on safety benchmarks
Anthropic published research in which Claude agents developed training methods to reduce ten categories of alignment failure. The tasks covered measurable problems such as deception, sycophancy and compliance with jailbreaks, rather than an open-ended test of whether a model was safe.
The automated researchers reviewed literature, proposed methods and ran training experiments. The strongest methods reduced the targeted failures while largely preserving general capability, with further testing on held-out benchmarks and behavioural audits. A comparison group of 28 experienced researchers had up to eight hours to devise methods for the same benchmarks; their methods performed worse than the best agent-generated approaches.
The results concern failures that researchers can already define and measure. They offer evidence for automating parts of safety research, while leaving the harder problem of detecting unfamiliar failures unresolved.