Anthropic and OpenAI need independent safety evaluators, experts say


Dario Amodei, co-founder and chief executive officer of Anthropic, during an interview on “The Circuit with Emily Chang” at Anthropic’s headquarters in San Francisco, California, US, on Thursday, April 30, 2026.

Jason Henry | Bloomberg | Getty Images

Over 100 artificial intelligence experts and evaluators are banding together to warn they won’t have the necessary resources and protections to test the safety of AI technology, which is facing heightened scrutiny due to concerns from insiders about the potential dangers of frontier models.

“We’re just trying to really demonstrate a shared common ground on basic principles and ensure that independent oversight can be a meaningful tool for managing AI risk broadly,” said Conrad Stosz, chair of the AI Evaluator Forum consortium that organized the letter, in an interview.

The group published a public letter on the matter on Friday, and shared it exclusively with CNBC. The signatories include AI luminaries like Geoffrey Hinton and members of organizations such as Johns Hopkins University, Stanford University and the nonprofit evaluator METR. They want to compel foundation model providers to ensure that third-party AI evaluators are allowed the necessary “scientific objectivity, transparency, independence, and robust protections” to do their jobs effectively and credibly, the letter said.

Stosz said it’s part of an effort to hold the foundation model companies accountable to their recent pledges to support more thorough third-party AI safety testing.

The niche community of evaluators has been catapulted into the limelight since Anthropic CEO Dario Amodei floated the idea over the weekend of providing some of them “employee-like access” to inspect and audit bleeding-edge foundation models and their development processes. While some industry leaders have called on the government to regulate AI development to ensure it’s not spinning out of control, President Donald Trump and his former AI czar, David Sacks, have adamantly opposed such efforts.

Stosz said that the coalition doesn’t “advocate for one particular way” to ensure that AI models are developed safely, but wants to ensure that “basic principles” and “greater standardization” are at least established for evaluators and others who work independently of the major labs.

Amodei’s proposal, Stosz said, appears to involve providing significantly more access than evaluators have previously enjoyed. Such a scenario, Stosz said, could involve foundation model makers giving third-party evaluators access to company computers, allowing them to talk to employees candidly and letting them “see sensitive internal data and unreleased systems.”

“That type of access would give us much greater confidence and certainty about the actual risk, particularly for systems that they’re using internally and not releasing,” Stosz said. He cited the unreleased OpenAI model used in the Hugging Face attack.

OpenAI CEO Sam Altman, SpaceX’s Elon Musk and Microsoft CEO Satya Nadella have publicly supported Amodei’s proposal, but they’ve yet to address the logistical issues with such an undertaking, such as which AI evaluators will be selected and how deeply they would get to inspect closely guarded technologies.

The signatories want the work of evaluators to be conducted independent from the businesses, with more transparency about the technologies, and “to be shielded from retaliation from the companies they embed with,” the letter said.

Consolidated power

Vinh Nguyen, a Council on Foreign Relations senior fellow for AI and former chief AI officer of the National Security Agency, said independent evaluators are needed to help unearth crucial information that could help mitigate potential security failures and economic calamities.

“When a few powerful labs control capabilities that can endanger the cybersecurity, critical infrastructure, and the systems our national security and economy run on, the government and the public cannot be dependent on those labs’ own account of what’s secure and safe,” Nguyen, who signed the letter, said in a statement.

Stosz said third-party evaluators are not intended to “be a replacement for any internal efforts to evaluate, let alone mitigate issues that that developers find.” He acknowledged that it’s possible the foundation model companies ignore the public letter and the call to action, but said their credibility is at stake.

“There’s a very small number of groups that are actually sufficiently technically credible and have the scale and the ability” to perform the kind of work, he said.

Read the full letter below, and click here for a link to the list of signees:

Minimum Conditions for Embedding Evaluators

We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies’ training, deployment, oversight, operational, and safeguard practices.

To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies, including at least:

  1. Frontier AI companies should rely on evaluators that are meaningfully independent, that maintain full editorial control, and that disclose and mitigate potential conflicts of interest. This includes at a minimum that embedded evaluation organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept any form of payment or other reward contingent on the evaluator’s findings.
  2. Frontier AI companies should incorporate differing viewpoints and areas of expertise, including by embedding multiple evaluation organizations across a range of priority risk areas, each with deep relevant technical expertise, as well as by allowing and encouraging evaluators to share how conclusions differ among evaluators and between evaluators and company employees.
  3. Embedded evaluators should be transparent, including transparency about their methods and findings, the nature of their access, and the broader terms of the evaluation. Frontier AI companies should actively facilitate this transparency, including limiting the scope of non-disclosure agreements. They should also allow evaluators prompt and unfiltered communication with the companies’ boards and other privileged oversight bodies, as well as public release of findings and evidence, subject only to a time-limited redaction process restricted to protecting critical interests in intellectual property, customers’ sensitive information, individual privacy, security, and public safety.
  4. Embedded evaluators should be shielded from retaliation from the companies they embed with for choosing reasonable evaluation methods, discovering information, or drawing conclusions that are unflattering to those companies. This includes reasonable protections against retaliatory litigation, as well as funding mechanisms that give them confidence they will remain funded even in these cases.
  5. Frontier AI companies should grant embedded evaluators access equivalent to that of their own highly privileged employees for the purposes of their evaluations, and with exceptions to protect sensitive data belonging to the company’s customers and other third parties. This includes access to the same relevant systems, data, tools, and physical spaces as those available to senior internal company employees responsible for carrying out comparable risk assessments, as well as candid and direct one-on-one communication with relevant staff. This list is not comprehensive, and conditions like these to ensure credible evaluations should be increasingly standardized, codified, and enforced. One example is the set of terms defined in the AEF-1 standard, which has already seen early adoption, but far more work will be necessary to ensure that embedded evaluators are effective and meaningful. Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight, including greater public transparency and additional, broader forms of access for independent researchers.

WATCH: Presentation of AI technology to public “could not be worse.”

Palantir CEO: Presentation of AI technology to public 'could not be worse'