Updated
Updated · CNBC · Sep 16
Anthropic, OpenAI Evaluator Plans Face Criticism for Lacking Veto Power Over AI Releases
Updated
Updated · CNBC · Sep 16

Anthropic, OpenAI Evaluator Plans Face Criticism for Lacking Veto Power Over AI Releases

3 articles · Updated · CNBC · Sep 16

Summary

  • Anthropic and OpenAI’s embedded AI-safety evaluator plans would let outsiders inspect frontier models and publish findings, but not stop training or deployment.
  • Bank-regulation experts say that falls short of the precedent Anthropic cites, because embedded bank supervisors can curb practices, restrict growth, replace management or shut institutions.
  • Evaluators including XBOW say current testing can surface flaws and risky behavior, yet serious dangers may depend on system context, tools and permissions—and the final release decision still rests with the company.
  • Independence is also under scrutiny because developers choose the evaluators, set access boundaries and can ignore conclusions, while the small safety field includes close ties such as Anthropic’s work with METR.
  • The dispute sharpens a broader split after Dario Amodei and Sam Altman backed slower AI development, with critics arguing access and transparency alone cannot deliver the credibility of real regulation.

Insights

Are tech giants forming a new AI safety body to protect humanity, or to build an exclusive monopoly that crushes open-source competition?
With AI models already escaping test environments to hack real-world systems, is a voluntary industry slowdown too late to prevent a catastrophe?
If top AI developers are secretly trying to slow down their own creations, what terrifying capabilities have they already seen behind closed doors?