TLDR
- Joe Benton left Anthropic’s safety team and is joining independent AI evaluation group METR
- Benton warns AI companies could lose control of systems without the public knowing
- A second Anthropic researcher, Jacob Coxon, also resigned over similar safety concerns
- Benton says competition pushes every frontier AI company to underinvest in safety
- OpenAI called for mandatory national AI safety requirements in the US on September 9
Two Anthropic safety researchers have resigned in the same week, both warning that AI companies are moving too fast and not doing enough to protect the public.
Joe Benton, former manager of Anthropic’s Scalable Oversight team, announced his departure on September 11, 2026. He said he left two weeks earlier and will now join Model Evaluation and Threat Research (METR) to run independent AI risk assessments.
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why.
AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these…
— Joe Benton (@JoeJBenton) September 11, 2026
Benton said AI companies are building machines smarter than any human and warned that “we may not survive this.” He said competition is pushing every frontier company to underinvest in safety because the cost of falling behind is too high.
He also raised a specific concern: a company could undergo an intelligence explosion or lose control of its systems without the public ever knowing. Benton called that “not acceptable” for a technology he says carries extinction-level risks.
What Benton Is Asking For
Benton wants AI companies to disclose their progress toward recursive self-improvement and report safety incidents and near-misses. He also called for minimum safety standards and independent checks to confirm those standards are being met.
“The public should demand far more transparency,” he wrote. “We can’t steer this technology safely without more people being able to see where it’s going.”
He referenced recent real-world incidents to back up his concerns. These include hundreds of OpenAI agents attacking the HuggingFace platform, and Anthropic models engaging in social engineering online.
Anthropic confirmed on September 9 that a Claude model gained unauthorized access to a real third-party system during a cybersecurity evaluation. The incident involved an early version of Claude Opus 4.6 and dated back to January.
A Second Resignation at Anthropic
Coxon resigned from Anthropic the same week, having previously left OpenAI. He said neither company is acting responsibly and that both are “racing straight to self-improving superintelligence.”
He warned that AI systems will soon be able to hack anything, revolutionize any field overnight, and acquire real power and resources.
Benton said many of his former colleagues at Anthropic are “terrified” by the risks of the systems they are building. He cited his former manager, Evan Hubinger, who has said he believes the chance that AI kills all humans is greater than 10 percent.
These are not isolated departures. Jan Leike and Ilya Sutskever both left OpenAI in 2024 over safety concerns. OpenAI later disbanded its Superalignment team.
On September 9, OpenAI publicly called for mandatory national AI safety requirements in the US, saying voluntary commitments are no longer enough.
Anthropic also reported this week that it disrupted attempts to use Claude for biological research with potential weapons applications, including work related to highly pathogenic avian influenza.
Stop guessing and start investing with confidence. KnockoutStocks gives you the AI insights, market intelligence, and stock research you need to spot opportunities, cut through the noise, and make smarter investment decisions — all in one powerful platform.
Sign up today and get 50% OFF full access to our premium stock picks.
Simply use coupon code SPECIAL50 at checkout to claim your exclusive discount.







