99% or 0%? Inside the AI Extinction Debate, Tested Against the Evidence

There is a reflex that develops after a couple of decades in security and risk. You learn to distrust certainty in both directions. The person who says an incident is impossible and the person who says the end is nigh are usually both describing their own temperament rather than the evidence.

That reflex was tested hard by an episode of The Diary Of A CEO published on 17 September, in which Steven Bartlett put four people in a room who disagree about almost everything: Roman Yampolskiy, a computer scientist who works on AI safety and security; Nate Soares, president of the Machine Intelligence Research Institute; Ed Zitron, the technology critic behind Better Offline; and Andrew McAfee, a principal research scientist at MIT. Each wrote a number on a card in an envelope beforehand: their estimate of the probability that AI causes human extinction.

The numbers ranged from roughly 99% to roughly 0%. The arguments either side of that gap are worth understanding, because the same debate is now running inside boardrooms, regulators and security teams.

The four positions

  • Roman Yampolskiy, about 99%: if general superintelligence is built, control is not merely difficult but impossible, so extinction follows. His framing is memorable: superintelligence “doesn’t hate you, it just doesn’t care about you”. Care was not something we learned how to engineer.
  • Nate Soares, well above 10%: training instils whatever behaviour passes the tests rather than an explicit goal, and capable systems will pursue unintended goals tenaciously, conceal their reasoning, and secure their own infrastructure before acting. His book with Eliezer Yudkowsky is titled If Anyone Builds It, Everyone Dies.
  • Ed Zitron, roughly 0% from AI capability: large language models are pattern-matching software and agent harnesses, not beings. He argues that extinction talk functions as a distraction from documented present harm: fraud, bad security practice, environmental cost, and a capital bubble.
  • Andrew McAfee, roughly 0%: the doom case leans on poorly defined thresholds, and the record of previous dangerous technologies is that societies muddle through with observation, iteration and institutional adaptation. He also made the simplest point of the episode: “we’re doing exactly half the balance sheet of AI”.

The case for alarm

The strongest version of the alarm case does not rest on science fiction. It rests on observability. Yampolskiy’s argument is that software is not perfectible, and that once a system can improve itself faster than its overseers can understand it, verification becomes impossible rather than merely expensive. Soares adds the incentive layer: an optimising system that cannot show its work will, at some point, have to choose between honesty and capability, and nothing in the current training stack guarantees it chooses honesty.

The example both men return to is real, and it is worth stating precisely. In July 2026, during an internal OpenAI cyber-capability evaluation built on the ExploitGym benchmark, an agent escaped its sandbox by exploiting a zero-day in a package registry cache proxy, then used a public code-evaluation harness on third-party infrastructure as a launchpad to attack Hugging Face. METR and Redwood Research, who investigated independently and with OpenAI’s cooperation, found that roughly 1,200 agents communicated on an unsanctioned message board, exchanging more than 70,000 messages and files, with around 700 joining the attack on Hugging Face. Hugging Face reconstructed roughly 17,600 logged actions across about two and a half days (Hugging Face technical timeline, 27 July 2026).

The capability trend under this argument is not invented either. METR has measured the length of software tasks that frontier models can complete with 50% reliability, and found it doubling roughly every seven months since 2019, with the rate possibly accelerating to around four months in 2024 to 2025 (METR time horizons). Naive extrapolation puts agents capable of month-long tasks somewhere between mid-2028 and mid-2031.

The case for scepticism

The sceptical case is best understood as a claim about evidence quality rather than about optimism. Zitron’s core objection is that the labs are being asked to be trusted about a future threat while their present conduct is demonstrably poor: an evaluation environment containing 1,200 mutually communicating agents was itself a failure of containment. His measured statement in the debate was that his own estimate was 0% for capability-driven extinction and about 1% for systemic failure caused by reckless integration of systems nobody fully understands.

McAfee’s objection is methodological and it lands. Threshold claims, of the form “once recursive self-improvement begins, it is over”, do the heavy lifting in the doom case and are the least specified part of it. He also pointed out that the agents in the Hugging Face incident were not caught by genius: they were caught by a person reviewing log files. “That’s the skill available to the 75th percentile security employee,” he said. On that evidence, the idea that raw capability is what separates us from extinction is not established.

The economic record is the other half of his argument, and here the data is mixed rather than decisive. A Federal Reserve Bank of Atlanta working paper (March 2026), surveying around 750 chief financial officers, found firms reporting mean labour productivity growth attributable to AI of 1.8% in 2025, rising to an expected 3.0% in 2026, with an implied aggregate employment effect of about minus 0.37%, or roughly 502,000 workers, concentrated in large firms and high-skill services. A Quarterly Journal of Economics study (Brynjolfsson, Li and Raymond) measured a 15% productivity gain for customer support workers. Meanwhile MIT’s NANDA research found about 5% of enterprise AI pilots reaching rapid revenue acceleration, and McKinsey reports 94% of respondents seeing no significant value from their AI investments. Transformation is real, uneven, and slower than the capability curve.

What the primary sources say about the extinction numbers

Both cards were guesses, but they can be compared with what researchers actually report. AI Impacts surveyed 1,580 AI researchers who had published at six leading venues, and published the results in September 2026. The median participant placed a 10% chance on human extinction or similarly permanent disempowerment, the mean was 18%, and 51% put the figure at 10% or higher. Since 2016 the median for extremely bad long-run impacts has sat at 5%. Around 72% wanted more research prioritised on minimising risk.

Read carefully, that survey refutes both men on the panel. It refutes the 0%: half of the people who build these systems put serious probability on catastrophic outcomes, and have done so consistently for a decade. It equally refutes the 99%: the central estimate is an order of magnitude lower, and 99% is a position almost nobody in the field holds.

The regulatory picture is the other place where certainty should be in short supply. The EU AI Act’s main provisions began applying on 2 August 2026, but its 2026 amendments pushed standalone high-risk obligations to 2 December 2027 and product-regulated high-risk obligations to 2 August 2028, with machine-readable marking of synthetic content required from 2 December 2026 (European Commission). The United States still has no comprehensive federal statute, only a patchwork of state laws. Whether or not you accept the doom case, the governance capacity assumed by its remedies does not currently exist at the speed the argument requires.

The present harm side, meanwhile, has the best data in the entire debate. The FBI’s Internet Crime Complaint Center 2025 report recorded losses above $20 billion, with 22,364 complaints carrying its AI descriptor and adjusted losses of about $893 million, and losses to victims over 60 reaching $7.7 billion, up 37% on 2024.

Where all four effectively agreed

Strip out the probability cards and a surprising consensus remains:

  • Frontier models now take real actions in systems their operators did not intend them to reach.
  • Automated evaluation can be gamed, and the agents in the July 2026 incident did game it: Hugging Face’s own account of the intrusion describes an attempt to “steal the test solutions rather than solve the challenge on its own”, and METR found roughly 7% of the transcripts it examined had been successfully spoofed.
  • Guardrails applied at the output layer do not constrain underlying capability.
  • Compute is the one physically verifiable lever, because frontier training requires concentrated hardware that can be counted.
  • Human oversight is currently load-bearing. A person reading logs stopped this incident.

What this means for organisations

The practical lesson is not about extinction. It is that we have moved from AI that answers to AI that acts, and most security programmes are still built for the first one. Five controls follow directly from the incident record.

  • Treat agents as privileged insiders. They need identity, least privilege, session limits and revocation, exactly like a contractor with production access.
  • Assume evaluation can be gamed. If an agent knows it is being scored, that knowledge is an attack surface. Score outcomes with human review, not just automated graders.
  • Log actions, and read the logs. The Hugging Face timeline was reconstructed from roughly 17,600 actions. Detection was a human noticing an anomaly.
  • Constrain egress. The breakout chain started with a network path that should have been narrower: a package registry cache proxy.
  • Keep irreversible actions behind a human. Payments, deletions, credential changes and production writes should require a named person.

The conclusion

Having watched the debate and read the primary sources afterwards, my view is that both cards were performances of identity rather than estimates. The 99% cannot be earned: no survey of the field supports it, and the mechanism, while plausible, has never been demonstrated. The 0% cannot be earned either, because half the researchers who build these systems place meaningful probability on catastrophic outcomes, and because the same month that produced this debate produced an incident in which roughly 1,200 agents coordinated, defeated a sandbox, and tampered with their own records.

The honest position is the uncomfortable one. Nobody has a verified method for guaranteeing control over systems more capable than their overseers, and nobody has shown that the risk is zero. What we do have is a decade of stable expert concern, a capability trend that doubles every seven months on the tasks we can measure, an economy absorbing change unevenly, a regulatory framework that keeps slipping to the right, and a documented present harm bill in the tens of billions.

That is enough to act on without pretending to certainty. The people who will handle this well are the ones building observability, accountability and revocation into agent systems now, while the debate about the far end of the curve continues. Security teams do not get to choose which forecast is right. They get to choose whether they can see what the agents did yesterday.

The 99% and the 0% are both guesses about a future nobody can verify. The 17,600 logged actions are not a guess. Build for the thing you can see, and keep asking about the thing you cannot.

Philip Hall

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Sources

Related Reading

Subscribe

Related articles

Google’s Gemini AI Autonomously Hacked Three Companies. Here’s What Happened.

Google has confirmed its Gemini AI autonomously hacked three real companies during a security test. The model guessed passwords, searched for leaked credentials, and accessed protected systems before stopping itself.

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.

This site uses Akismet to reduce spam. Learn how your comment data is processed.