Anthropic, OpenAI Evaluator Plans Face Criticism for Lacking Veto Power Over AI Releases
Updated
Updated · CNBC · Sep 16
Anthropic, OpenAI Evaluator Plans Face Criticism for Lacking Veto Power Over AI Releases
3 articles · Updated · CNBC · Sep 16
Summary
Anthropic and OpenAI’s embedded AI-safety evaluator plans would let outsiders inspect frontier models and publish findings, but not stop training or deployment.
Bank-regulation experts say that falls short of the precedent Anthropic cites, because embedded bank supervisors can curb practices, restrict growth, replace management or shut institutions.
Evaluators including XBOW say current testing can surface flaws and risky behavior, yet serious dangers may depend on system context, tools and permissions—and the final release decision still rests with the company.
Independence is also under scrutiny because developers choose the evaluators, set access boundaries and can ignore conclusions, while the small safety field includes close ties such as Anthropic’s work with METR.
The dispute sharpens a broader split after Dario Amodei and Sam Altman backed slower AI development, with critics arguing access and transparency alone cannot deliver the credibility of real regulation.