Former Anthropic researcher
He left
the lab.
Then sounded
the alarm.
Jacob Coxon resigned from Anthropic and took a stark warning about frontier AI into public view. What he said, what remains contested, and what the record actually shows.
What happened—and what did not.
On September 8, Jacob Coxon announced that he was leaving Anthropic after research roles at Anthropic and OpenAI. In a public thread, he argued that the leading labs were racing toward self-improving AI without an adequate safety plan.
The distinction matters: Coxon resigned. Available reporting does not say Anthropic fired him. News coverage has often called him a former researcher or departing employee; “whistleblower” is a descriptive label used by some commentators, not a formal legal status established by the available record.
AP, WIRED, CNN and Coxon’s own account consistently describe a voluntary departure.
The verifiable television interview used here aired on CNN’s Anderson Cooper 360. CNBC separately covered the story.
Coxon told WIRED that he had worked on pretraining research and described internal conversations in which colleagues used phrases such as “crunch time.” WIRED also reported his view that Anthropic was more safety-conscious than OpenAI, while he argued that competitive pressure could still force future trade-offs.
Watch the argument unfold.
CNN’s interview separates the near-term state of current models from Coxon’s concern about a future, recursively improving system.
“There is a very real possibility” of catastrophic outcomes.
Coxon’s claim is about a future capability threshold. In the CNN interview, he explicitly said current models do not pose an extinction risk and focused instead on the pace of improvement.
He pointed to advances in coding, mathematics and cyber operations as evidence that capability growth deserves close attention.
The core concern is whether increasingly autonomous systems can reliably remain under human direction.
Labs may want safeguards while fearing that a rival will advance faster, creating pressure against unilateral restraint.
Claims, corroboration and uncertainty.
One current Anthropic alignment lead, Evan Hubinger, publicly agreed with the broad concern and assigned a greater-than-10-percent personal probability to AI killing all humans within a decade, according to CNN and WIRED. That is a researcher’s risk estimate—not a measured probability or scientific consensus.
AP reported that two current Anthropic employees responded in agreement. The resonance inside the field is newsworthy, but it does not settle the technical debate. Forecasts about systems that do not yet exist depend on uncertain assumptions about capability growth, autonomy, access to tools, defensive safeguards and governance.
Coxon cited cyber incidents involving AI agents as warning signs. Anthropic’s own research page lists alignment, interpretability and frontier red-teaming programs, including work on cybersecurity incidents. Those programs document active safety research; they do not prove either that catastrophe is imminent or that the problem is solved.
What would a guarded future require?
In its response reported by CNN and WIRED, Anthropic emphasized both the benefits and unprecedented risks of AI, its testing of dangerous capabilities, and its support for lawful, verifiable coordination on the pace of powerful model releases.
Coxon proposed an initial agreement between major Western labs to avoid immediately pursuing recursive self-improvement, followed by broader international coordination. That prescription raises difficult questions: how capability thresholds would be defined, how compliance would be verified, how open research would be treated, and how states could coordinate without freezing useful safety work.
Explore safer AI workflows in Roseram →Test dangerous capabilities before release, with independent scrutiny and published methods.
Limit permissions, isolate sensitive tools, preserve human approval and support rollback.
Report failures consistently enough for researchers and regulators to learn from them.
Turn voluntary promises into rules that can be observed, audited and enforced.
The warning became two stories.
A researcher gave up a prestigious role to argue that competitive incentives were outrunning safety planning.
A dramatic extinction claim became a political and cultural symbol, often stripped of its uncertainty and technical conditions.
Read the underlying record.
This article was independently assembled from a televised transcript, Coxon’s published interview, wire reporting, secondary coverage, and Anthropic’s primary research materials. Roseram did not independently validate Coxon’s private conversations or calculate an extinction probability.