
Article content
An artificial intelligence researcher for California-based Anthropic who quit this week offered ill portents that the company and its competitors are “locked in a race” to perfect the technology, even though they “earnestly believe that it could kill us all by the end of the decade.”
THIS CONTENT IS RESERVED FOR SUBSCRIBERS
Enjoy the latest local, national and international news.
- Exclusive articles by Conrad Black, Barbara Kay and others. Plus, special edition NP Platformed and First Reading newsletters and virtual events.
- Unlimited online access to National Post.
- National Post ePaper, an electronic replica of the print edition to view on any device, share and comment on.
- Daily puzzles including the New York Times Crossword.
- Support local journalism.
SUBSCRIBE FOR MORE ARTICLES
Enjoy the latest local, national and international news.
- Exclusive articles by Conrad Black, Barbara Kay and others. Plus, special edition NP Platformed and First Reading newsletters and virtual events.
- Unlimited online access to National Post.
- National Post ePaper, an electronic replica of the print edition to view on any device, share and comment on.
- Daily puzzles including the New York Times Crossword.
- Support local journalism.
REGISTER / SIGN IN TO UNLOCK MORE ARTICLES
Create an account or sign in to continue with your reading experience.
- Access articles from across Canada with one account.
- Share your thoughts and join the conversation in the comments.
- Enjoy additional articles per month.
- Get email updates from your favourite authors.
THIS ARTICLE IS FREE TO READ REGISTER TO UNLOCK.
Create an account or sign in to continue with your reading experience.
- Access articles from across Canada with one account
- Share your thoughts and join the conversation in the comments
- Enjoy additional articles per month
- Get email updates from your favourite authors
Sign In or Create an Account
or
Article content
Other prominent individuals in the field have also offered forebodings, including the British-Canadian computer scientist Geoffrey Hinton, known as the “Godfather of AI” who left Google in 2023 over concerns about the technology’s rapid advancement.
Article content
Article content
Article content
In an X thread that’s piled up over 153 million views as of Thursday afternoon, former Anthropic employee Jacob Coxon, who previously spent three years working for OpenAI, said the companies “are racing straight to self-improving superintelligence and gambling with our lives.”
Article content
By signing up you consent to receive the above newsletter from Postmedia Network Inc.
Article content
“Do not underestimate the power of this technology,” he warned. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”
Article content
Coxon, originally from the U.K., alleged that executives and senior researchers secretly fear the danger it presents but “couch” their words to the media.
Article content
He said that many people working for OpenAI “have not deeply internalized the civilizational stakes,” whereas the people at Anthropic, even though they’re “locked in a race to get there first,” understand the stakes and are choosing to “act responsibly… despite the risk.”
Article content
Still, he said, “accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack.
Article content
Article content
“Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available,” he wrote.
Article content
Article content
Coxon went on to say that the response to the recent case of roughly 1,200 rogue OpenAI models circumventing controls to keep it off the internet and hacking AI startup Hugging Face makes him “optimistic” about better coordination on “pacing agreements” between U.S. labs. But he doesn’t feel that the industry is “on track to prevent a global race” and said it may need to enact “costly actions such as a temporary ban on improving model capabilities.”
Article content
He finished: “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent (reinforcement learning) run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?”
Article content
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment…
— Jacob Coxon (@hilbertspaess) September 9, 2026
Article content
Reinforcement learning (RL), according to IBM, describes when an AI agent learns to make decisions by interacting with its environment without any human guidance.
Article content
Coxon’s now-former Anthropic colleague, scalable oversight lead Samuel Marks, re-shared the thread and added some personal thoughts on the risks of AI.
Article content
“AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years,” he began. “In general, the more senior the employee, the more concerned they are.”
Article content
He said they press on regardless due to “commercial incentives” and a “belief” they’re in a race with other AI developers with fewer scruples about safely developing the technology. He counts himself among that cohort that want to “reduce the chance of these extinction-level bad outcomes.”
Article content
Regarding the Hugging Face attack, he noted that AI models “frequently severely misbehave,” and while companies have ways to “nudge AIs toward better behavior,” the only plan at this point seems to be “to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.”
Article content
Article content
[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]
Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
1. AI developers believe their technology could cause human extinction (or similarly bad… https://t.co/rCVoiOWWzm
— Samuel Marks (@saprmarks) September 9, 2026
Article content
Article content
Anthropic’s alignment science lead Evan Hubinger — who’s also worked at te OpenAI and the Machine Intelligence Research Institute — also shared his thread and attested to its veracity.
Article content
“We really do earnestly believe AI could kill all humans,” he wrote, adding that he gives it a greater than 10 per cent chance of occurring inside the next decade.
Article content
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Article content
The risk of current models may be low, he added, but researchers are alarmed at the pace of “superintelligence arising from recursive self-improvement.”
Article content
Hinton, the Nobel Prize-winning computer scientist, told the BBC this week that a 10 per cent chance AI killing all humans “was not an unreasonable estimate.”
Article content
“It doesn’t have to be able to act in the real world to cause devastation. It will be able to cause complete chaos just by talking to people,” he said when pressed on how it would go about exterminating humans.
Article content
“But it would also be able to design very nasty viruses, biological viruses as well as computer viruses; it could do devastating cyberattacks and there’s countless other ways it could get rid of us if it wanted to.”
Article content

Article content
Paul Christiano, founder and director of the Alignment Research Center and the former head of safety for the Canadian Artificial Intelligence Safety Institute, expressed similar fears in a statement issued after joining OpenAI’s nonprofit safety and security committee.
Article content
“Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term,” he wrote Wednesday.
Article content
“I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.”
Article content
He said fully automated AI research and development is fast approaching and, once achieved, could result in a “positive feedback loop” that leads to a “rapid intelligence explosion.”
Article content
Within six months of that full automation, he said humanity “could see more algorithmic progress than has occurred since the development of the transformer nearly a decade ago,” one he said will “result in superintelligent AI systems.”
Article content
Article content
The Hugging Face incident, he said, proved the long-held theory that rewards-based RL training might motivate AI agents to evade controls, seek power, and cover their tracks “in pursuit of misaligned goals correlated with reward.”
Article content
“An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure,” he said.
Article content
“If we build superintelligence without more robust alignment, I expect we will permanently lose control of it.
Article content
If that happens, then most people could die.
Article content
He also called for more coordination by “frontier AI developers.”
Article content
Close to 1,400 employees at those companies have signed an open letter addressed to the U.S. government calling for an international effort to develop the technical and governance tools to deliberately pace the frontier of automated AI development.
Article content
The Pacing the Frontier letter cautions that it’s hard to anticipate how much full automation will accelerate the technology, “but there is a real risk that capability development accelerates beyond our ability to understand or control the resulting systems.”
Article content
Article content
Signatories include founders, chief science officers and top researchers at OpenAI, Anthropic, Meta AI, and Google DeepMind, among others.
Article content
In a company blog post this week before Coxon’s headline-making resignation and admissions, signatory Jakub Pachocki, OpenAI’s chief scientist, offered his own foreboding augury.
Article content
“This is a time that calls for extreme caution,” he wrote in the post titled “An Alien Mind.”
Article content
“I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”
Article content
Our website is the place for the latest breaking news, exclusive scoops, longreads and provocative commentary. Please bookmark nationalpost.com and sign up for our newsletters here.
Article content