Anthropic Pledges Employee-Level Access for Independent AI Evaluators
Dario Amodei's plan to slow AI development starts with outside reviewers inside Anthropic's offices, a commitment it acted on six days later by hiring Accenture.
Estimated reading time: 7 minutes
TL;DR
- Anthropic’s chief executive, Dario Amodei, published an essay arguing the AI industry should deliberately slow how fast it makes models more capable, so safety checks can keep up.
- His first move: promise outside reviewers the same kind of access Anthropic’s own safety staff get, plus the right to publish what they find without company edits.
- Six days later, Anthropic named the Accenture division Faculty as that first outside reviewer, with each company planning to put in at least $1 billion over five years.
- OpenAI’s Sam Altman said he would match the access pledge, but Meta’s Mark Zuckerberg and Nvidia’s Jensen Huang publicly rejected the idea of an industry-wide slowdown.
- A critic writing in Tech Policy Press warns that without government enforcement, the plan amounts to the industry policing itself.
What happened
On September 12, 2026, Amodei published an essay titled “We Must Pace the Frontier” on his personal website, arguing that AI companies should deliberately slow the rate at which model capabilities improve so that safety work can keep pace — while stressing this does not mean halting model training or technical progress altogether, according to the essay.
As the first, unilateral step of that plan, Amodei committed Anthropic to giving outside evaluators permanent, employee-like access to its offices and systems — permissions mostly comparable to what Anthropic’s own internal risk-assessment teams already have — so they can check the company follows its safety practices, report on incidents, and assess how models behave during training, he wrote. He described the arrangement plainly: “Desks in our offices, access badges, and company laptops. Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have.” Those evaluators, he added, should also be free to publish what they find — including about the access they were or weren’t given — without Anthropic editing it first, aside from limited redactions: “External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive—without editorial control by Anthropic.”
A second step Amodei proposed is voluntary coordination among AI companies based in democratic countries to set shared safety standards and limits on how fast capabilities advance. Because that kind of coordination between competitors raises antitrust concerns, he said it would help if the U.S. government mediated or enabled the talks by issuing a narrow waiver for safety-related conversations, and he floated — as one possible, not finalized, mechanism — a system where models that reach a given capability level would need certifications of specific safety properties attached, according to the essay. A third, international step sketches four escalating levels of possible agreement between governments, from banning narrow and obviously dangerous uses such as AI-assisted bioweapons production, up to mandatory pre-release safety testing, a limit on how fast AI systems could improve themselves, and finally a full slowdown of AI development overall — which Amodei says is unlikely soon given the difficulty of verifying compliance and catching countries that cheat, per the essay.
Six days later, on September 18, Anthropic followed through: it announced it had selected Accenture’s AI division, Faculty, as its first embedded evaluator, calling the move a step toward the commitment in Amodei’s essay. Per the announcement, Faculty’s evaluators will work inside Anthropic with employee-comparable access, watching models take shape during training, following the decisions that govern how models are built and deployed, and speaking directly with staff, while conducting red-teaming (deliberately probing a system for weaknesses), alignment assessments and tests of model safeguards. Anthropic and Accenture each expect to invest at least $1 billion over the next five years in building this evaluation capacity, with Anthropic funding Accenture’s work directly. TechCrunch independently confirmed the arrangement and reported it is non-exclusive: Anthropic said it is also in discussions with the research nonprofit METR and other potential evaluators.
Reaction from other companies has split. OpenAI’s chief executive, Sam Altman, wrote on X on September 12 that pacing “has been a primary topic” of internal OpenAI discussions in recent weeks: “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.” Meta’s chief executive, Mark Zuckerberg, publicly rejected the case for an industry-wide coordinated slowdown, arguing market pressure already pushes labs toward safer systems: “People won’t want to use agents that are misaligned with them and that don’t do what they ask, so labs have a strong natural incentive to make their models more aligned.” Nvidia’s chief executive, Jensen Huang, in a CNBC interview reported by Competition Policy International and published on PYMNTS, dismissed calls for new antitrust waivers or laws to enable a coordinated slowdown: “The fact that we need new laws, new antitrust laws, or new regulations, so that these companies could do their fundamental engineering and do it properly before they release products, that is just completely unnecessary.”
What this means (and what it does not)
Amodei’s essay turned a general safety argument into a specific, checkable commitment: a defined level of office and system access, and a public promise about who controls publication of findings. Anthropic acting within six days to name a paid, chosen evaluator shows the company treating that pledge as something to implement immediately rather than leave as a proposal for others to debate, as shown by its own announcement and by TechCrunch’s confirmation. Altman’s public agreement, even without operational details, shows the idea is not being dismissed by OpenAI, Anthropic’s closest large-model rival.
What it does not mean is that the industry has reached consensus. Zuckerberg and Huang each publicly rejected the case for a coordinated slowdown or antitrust accommodation, arguing existing market and engineering incentives are already sufficient. Nor is this a binding international agreement: the four-level framework is Amodei’s own proposal, not something any government has adopted, and it is not a pause in AI development — Amodei explicitly said pacing does not mean halting training. TechCrunch reported that some critics of Anthropic’s safety approach view the embedded-evaluator scheme as effectively letting the industry police itself and avoid stronger, externally imposed accountability. Writing in Tech Policy Press, Dave Karpf makes a similar argument: “It is a mistake to let the AI industry shape the contours of its own regulatory system, even if we grant the good intentions of its leaders.” He warns that without government enforcement, “Third-party, independent evaluators will either become a revolving door to the industry, or else they’ll be utterly outmatched.”
What we still do not know
No source describes a measurable capability threshold, enforcement mechanism, or penalty that would apply if a company that publicly commits to pacing or to embedded evaluators fails to follow through — Amodei’s own essay presents the capability-linked checkpoint idea only as one possible, unfinalized scheme. Because Anthropic selects, scopes and funds its own evaluators — Accenture’s Faculty, and reportedly METR in discussion — neither Anthropic’s announcement nor TechCrunch’s reporting addresses how a conflict between evaluator independence and being a paid partner would be managed. No source names an existing international body — a UN agency or treaty organization, for example — that would administer the four-level framework Amodei describes; it lays out levels of possible agreement, not a proposed administering body. Whether any government has agreed to mediate the antitrust concerns Amodei raises, or to issue the narrow safety-conversation waiver he calls for, is not addressed by any source consulted. OpenAI’s scope, timeline and choice of evaluator organizations beyond Altman’s initial post have not been detailed elsewhere. Backing for the plan is not industry-wide: alongside Altman’s endorsement, Zuckerberg and Huang have publicly rejected the case for a coordinated slowdown or antitrust accommodation. And because the Accenture/Faculty arrangement was only days old as of the most recent reporting available, with no public evaluator findings yet published, whether it will function independently of Anthropic in practice remains untested.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://darioamodei.com/post/we-must-pace-the-frontier — Dario Amodei
- https://www.anthropic.com/news/accenture-embedded-evaluation
- https://techcrunch.com/2026/09/18/anthropics-first-embedded-evaluator-is-accenture/ — Tim Fernholz
- https://x.com/sama/status/2098811563415150910 — Sam Altman
- https://fortune.com/2026/09/16/mark-zuckerberg-meta-ai-safety-jensen-huang-dario-amodei/ — Marco Quiroz-Gutierrez
- https://www.pymnts.com/cpi-posts/nvidias-huang-rejects-ai-slowdown-plan-and-call-for-antitrust-waivers/ — CPI
- https://www.techpolicy.press/who-should-pace-the-frontier-not-dario-amodei/ — Dave Karpf
Editorial check, counted automatically
- 7 sources cited
- 28 inline-linked claims
- 0 unsourced claims found
- 0 banned words found
- 2 numbers without context
Also available in Portugues (BR)