Anthropic commits to embedded external safety evaluators

Anthropic chief executive Dario Amodei proposed a three-stage approach to slowing frontier AI development while safety work catches up. The first step is a unilateral commitment by Anthropic to embed external evaluators with ongoing access comparable to its internal risk teams.

The planned reviewers would be able to assess training processes, investigate incidents and publish findings without Anthropic exercising editorial control. Narrow redactions would remain possible for confidential or security-sensitive information, and reviewers could disclose when a redaction affected their conclusions.

Further steps call for coordination among companies in democratic countries and, eventually, international agreements. Those remain proposals rather than completed arrangements. The essay says an external review team will be invited in the near future.

The commitment moves scrutiny closer to the work itself. Its value will depend on what reviewers can actually inspect and report, including when their findings are unfavourable.

Anthropic commits to embedded external safety evaluators | AI | Blog