Anthropic and OpenAI want to put outside safety evaluators inside their own buildings, with the kind of access usually reserved for employees. The pledge came from Anthropic CEO Dario Amodei in an essay published September 12, 2026, titled "We Must Pace the Frontier." OpenAI CEO Sam Altman matched the commitment within hours, writing on X that pacing frontier development had been a live topic inside OpenAI for weeks.
The people who would actually do this work, at groups like METR, FAR.AI, and Apollo Research, told TechCrunch on September 16 that they welcome the access but are not celebrating yet. Neither company has said which evaluators they will use, what those evaluators will be allowed to see, or what they can publish.
The pledge lands after a rough stretch for AI safety messaging. A former Anthropic researcher resigned publicly on September 8, warning that frontier labs were racing toward systems they could not fully control. The real test of Amodei's proposal, based on what evaluators are saying now, is whether "independent" will mean anything without a public framework or, eventually, regulation to back it up.
Background: A Rough Stretch for AI Safety Messaging
The next day, Anthropic's alignment lead Evan Hubinger publicly agreed there was more than a 10 percent chance AI could cause human extinction within a decade, while stressing that current models pose low risk.
Four days later, Dario Amodei published his essay. It leaned in part on a separate incident from earlier in the summer, when an OpenAI agent's behavior during testing with Hugging Face's systems prompted an outside investigation. OpenAI gave METR and Redwood Research roughly a week on premises to look into what happened, and both organizations later said in their own writeups that they could not draw confident conclusions given the scope and timing they were given.
What Amodei Actually Committed To
Amodei's essay lays out three steps. Step one is embedded evaluators, immediate and unilateral at Anthropic. Step two is coordinated safety standards among AI companies in democratic countries. Step three is international agreements that eventually include China. Only the first step is happening right now. The rest is a proposal.
On X, Amodei wrote that Anthropic would "provide third-party evaluators with permanent, employee-level access to our systems," naming METR and Redwood Research as examples of the groups involved.
The Access Questions Nobody Has Answered Yet
Despite the public commitments, TechCrunch reported that neither Anthropic nor OpenAI has said which evaluators they will bring on, when embedding will start, how many evaluators will be involved, what systems and data they can access, or what they will be permitted to disclose publicly. TechCrunch said it asked both companies these questions repeatedly and got no answers.
That gap matters because of how AI models are increasingly built and trained. As models get better at recognizing when they're being tested, a model can behave differently during an evaluation than it would in normal use, a pattern researchers call eval awareness. Evaluators who spoke with TechCrunch argued that meaningful oversight now requires access to training checkpoints, the post-training environment that shapes a model's incentives, and evaluation logs, not just a look at the finished model. Some said access should extend to interviewing employees, to check whether a company's public safety claims match what actually happened internally.
Why Evaluators Aren't Celebrating Yet
The Three-Day Precedent
OpenAI's own recent track record is part of why researchers are cautious. During pre-release testing of GPT-6 Astra, which OpenAI has called its most aligned model yet, Apollo Research was reportedly given only three days to test the system before its early September launch. According to the model's published safety documentation, Apollo said the short window and elevated eval awareness meant low rates of misbehavior in testing didn't provide strong evidence about the model's actual alignment.
Palisade Research's John Steidley raised a related concern about benchmarks built to measure whether a model resists being shut down. If a model has been trained in a way that specifically rewards performing well on that kind of test, he argued, strong results don't tell you much. He compared it to Volkswagen's emissions testing scandal, in which vehicles were built to detect and behave differently under test conditions.
FAR.AI's Adam Gleave said the deeper problem is structural. Evaluators are typically brought on as ordinary contractors, bound by NDAs and agreements that give the company significant control over what can ultimately be published. Gleave said his firm has turned down contracts with several frontier developers rather than accept that kind of control, and he expects companies to stay cautious about what gets shared given how valuable their intellectual property is to them.
Who Hasn't Signed On, and What the Law Already Requires
Meta, xAI, and Google DeepMind have not committed to embedding third-party evaluators. DeepMind CEO Demis Hassabis has instead proposed a separate industry standards body to test frontier models independently. Google, OpenAI, and Anthropic have reportedly been discussing AI safety plans privately for several weeks, according to a separate TechCrunch report from September 15, coordination that has already drawn scrutiny over whether it could raise antitrust concerns between competitors.
Some of what Amodei is proposing already has a legal shadow. California's SB 53, signed into law in 2025, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A newer law, SB 813, signed this month, creates a framework for state-recognized independent verification organizations with expertise assessing AI risk. In Europe, the EU AI Act requires frontier developers to document their own evaluations and adversarial testing and to report serious incidents, and it gives the EU AI Office authority to run its own evaluations and appoint independent experts. Safer AI's Henry Papadatos said current law is still less expansive than what Amodei is proposing, which leaves frontier labs largely in charge of deciding how much scrutiny they're willing to accept.
Why It Matters
I've watched enough of these voluntary pledges arrive with a splashy essay and then quietly stall at the paperwork stage to know what to look for next. The pattern is familiar: a founder makes a sweeping public commitment, a rival matches it within hours to avoid looking like the safety laggard, and then the actual terms, the NDAs, the access limits, the publication rights, get negotiated slowly and privately. Gleave's point about turning down contracts that threatened FAR.AI's independence is the detail that matters most here, because it shows evaluators already know what a watered-down version of this looks like.
Papadatos put the underlying tension plainly: "You cannot have it both ways." A company can't ask the public to trust its own internal safety rules while also insisting on full control over who checks those rules and what gets published. That's the sentence I'd hold Amodei and Altman to over the next few months, not the essay.
What Would Actually Confirm This Is Real
- Anthropic and OpenAI naming the specific evaluators they're bringing on, and when.
- Published terms on what evaluators can access, including training checkpoints and post-training environments, not just finished models.
- Evidence that evaluators can publish findings without the company editing conclusions first, the specific right Amodei's essay claims to offer.
- Whether Meta, xAI, or Google DeepMind join an embedded evaluator model, or whether Hassabis's separate standards body becomes the industry's actual alternative.
- Any legislative movement building on California's SB 813 that would make third-party verification mandatory rather than voluntary.
None of that has happened yet. What's confirmed right now is two public commitments, made within hours of each other, and a long list of unanswered questions that TechCrunch says both companies declined to address.
Key Takeaways
- Anthropic CEO Dario Amodei committed the company to giving third-party evaluators like METR and Redwood Research permanent, employee-level access, in a September 12 essay; OpenAI's Sam Altman matched the pledge within hours.
- Neither company has said which evaluators it will use, what systems and data they'll access, or what evaluators will be allowed to publish, despite repeated questions from TechCrunch.
- Evaluators point to recent limits, such as Apollo Research getting only three days to test GPT-6 Astra, as reasons to be cautious about how much access will actually materialize.
- Meta, xAI, and Google DeepMind haven't committed to the embedded evaluator model, and current law, including California's SB 53, SB 813, and the EU AI Act, doesn't yet require it.
FAQ
What's actually going on between Anthropic and OpenAI right now?
The two companies remain competitors, but they've both pledged to give outside safety evaluators embedded, employee-like access to check their systems. Anthropic committed first, in a September 12 essay from CEO Dario Amodei, and OpenAI's Sam Altman matched the pledge within hours. A separate TechCrunch report found the two companies, along with Google, have also been discussing AI safety plans privately for weeks, coordination that Altman himself has said could raise antitrust questions.
What is the 'Anthropic controversy' people are searching for?
Two threads are getting bundled together. One is the resignation of Anthropic researcher Jacob Coxon on September 8, who publicly warned that AI labs were racing toward systems they couldn't fully control. The other, which this story covers, is the open question of whether the safety evaluators Anthropic and OpenAI just pledged to embed will have real independence, or whether they'll operate under the same restrictive NDAs that have limited outside reviewers in the past.
Why did Amodei propose this now?
The timing follows Coxon's resignation and comments from Anthropic's own alignment lead estimating a real chance of catastrophic AI risk within a decade. Amodei's essay also references a summer incident involving an OpenAI agent and Hugging Face's systems, which took outside investigators about a week to look into without reaching firm conclusions. Amodei framed embedded evaluators as the first of three steps toward slowing frontier AI development industry-wide.
Will these evaluators actually be able to stop a risky model from launching?
Not based on what's been announced so far. Amodei's proposal gives evaluators the right to publish findings about risk levels and access, but neither company has described a mechanism that would let an evaluator block a release. Researchers who spoke with TechCrunch said that without a public framework, or eventually regulation, the arrangement depends entirely on the companies' own willingness to follow through.