AI experts call for protections and independence for embedded model safety evaluators
|
The Facts
- More than 100 AI experts and evaluators signed a public letter on third-party model evaluation.
- The letter addressed evaluations of frontier AI systems at companies including Anthropic and OpenAI.
- Signatories called for evaluators to have resources and protections needed to assess AI-model risks.
- The letter called for transparency conditions, including limits on the scope of nondisclosure agreements.
- Anthropic and OpenAI leaders committed to allowing third-party evaluators employee-level access.
- Anthropic selected Accenture to embed evaluators to test its frontier AI models.
- Anthropic and Accenture each expect to invest at least $1 billion over five years in evaluation capacity.
Context
What are embedded AI evaluators?
They are outside evaluators working within an AI company with access similar to employees, allowing them to observe model development, safety decisions and deployment processes. Investing.com TechCrunch
What would Accenture’s team do at Anthropic?
Accenture’s Faculty unit is expected to evaluate and red-team Anthropic models, conduct alignment assessments and test model safeguards. Investing.com TechCrunch Asian News Internat…
Why are signatories seeking additional conditions?
They argue that evaluators need independence, protections, resources and transparency to examine rapidly advancing AI capabilities and risks effectively. CNBC Business Insider Verge
Where Left and Right agree, and where they split
Left and right largely agree on this one.
- Where Left and Right agree
- Self-selected, company-funded evaluators like Accenture cannot substitute for independent oversight, so evaluators need guaranteed resources, protections, and limited NDAs.
- Where Left and Right differ in emphasis
- Both reject the Accenture arrangement as sufficient — one because it crowds out smaller independent monitors and privatizes findings, the other because it lacks any enforcement mechanism.
- Why they won’t converge
- The disagreement is about trust in institutions: whether a self-selected, self-funded evaluator arrangement can ever substitute for structural, enforceable independence, a question the facts of access and funding don't resolve.
How left and right read it
When the examined company also picks and pays its examiner, safety becomes a contract term rather than a public right. Scale is not independence. Anthropic selected Accenture, with each side expecting at least $1 billion over five years — a choice that routes around the smaller monitors already doing this work, which is why more than 100 signatories demanded resources, protections and limits on nondisclosure: findings about frontier risk belong to the public.
“In picking Accenture, a consulting behemoth that operates in more than 100 countries, Anthropic has initially bypassed Silicon Valley-based nonprofit monitoring services that critics perceive as being too close to the company.” — Washington Post
A promise made by the party under examination is not a check on it, because independence has to be structural rather than volunteered. Anthropic and OpenAI leaders committed to giving third-party evaluators employee-level access, yet Anthropic itself selected Accenture to embed the testers — which is why the more than 100 signatories press for real resources, protections, and limits on the scope of nondisclosure. Who enforces the commitment?
The receipts — all 45 sources
Wire services (5)
Independent coverage (40)
Facts first. Then every angle.
The day’s biggest stories in one short brief — the facts everyone agrees on, then the competing values behind the headlines. Free in your inbox.