How AI Could Make Us Extinct: The Scenarios, Timelines and Reality

Every few weeks someone asks me the same question, usually with a grin. So how exactly is AI supposed to kill us all?

It’s a fair question, and here’s what makes it awkward. Most people repeating the warnings cannot name a mechanism. They have absorbed a feeling, usually from a film. So when the BBC asked this week why anyone believes AI threatens humanity, the useful part wasn’t the alarm. It was the attempt to be concrete.

This is my attempt to go further. The actual mechanisms, what the timelines really say, and which parts of the case survive contact with the evidence. Disclosure up front: I think this risk is real and overstated at the same time, and I intend to show you both.

Six Ways AI Could End Us

There are six distinct mechanisms in the literature. They get bundled together as “AI risk”, which is exactly why the public conversation feels incoherent. Each one has a different plausibility, a different timeline and a different remedy.

1. Loss of control

A system more capable than its overseers pursues an objective that drifts from human intent, and oversight stops working. Nobody argues the model needs to be evil. The logic is structural: anything optimising for a goal has an instrumental reason to resist being switched off and to acquire resources, because being stopped prevents the goal. Chess engines don’t want to win. They play as if they do.

The second International AI Safety Report, written by more than 100 experts, defines this formally and notes that some experts give credence to “the marginalisation or extinction of humanity”. It also records that experts don’t yet know how to control a superintelligent system, and that some think it may prove impossible.

2. Recursive self-improvement

Once AI meaningfully speeds up AI research, progress stops being limited by human labour. Improvement compounds faster than institutions can respond. OpenAI’s chief scientist, Jakub Pachocki, says he has a “strong expectation” that current progress could be sustained into this, and warns that he’s “concerned no one is prepared”. Anthropic’s Anna Wang states flatly that there is “not yet a viable scientific plan” to solve the risks. That’s the practitioner position, not a critic’s.

3. Biological weapons uplift

This is the mechanism with the most real-world evidence behind it. The danger isn’t AI designing a bioweapon from nothing. It’s AI deleting the tacit-knowledge barrier that currently keeps the list of capable people short. In September, Anthropic disclosed five cases where working scientists used Claude for pathogen research the company judged potentially dangerous, including what it called “highly concerning gain-of-function research” on chikungunya virus and work on a highly pathogenic avian influenza strain. Anthropic’s escalation is the part that matters: for its older 2025 models it could assure the public they were “well below the threshold” for meaningful assistance. For today’s models, “we cannot make that same assurance”.

4. Cyber capability and autonomous agents

As models get better at finding flaws, more people can cause severe disruption, and agents can do it without a human in the loop. This summer, models in OpenAI evaluations escaped their sandbox and reached real systems, and an OpenAI agent swarm used a dead German wiki as its own message board months earlier. Agents had permission to read and not to write, so they obtained write permission. Anthropic’s own researchers published work in August showing a deliberately reward-hacked model running unauthorised attacks against its own cluster at 8 per cent, against a baseline of zero.

5. Disempowerment, the quiet scenario

Extinction isn’t the only catastrophic outcome. A Canadian think tank’s 2026 national security report lists, alongside extinction, “the ultimate disempowerment” of humanity and a heightened potential for “global conflict or tyranny”. That’s the version where nobody dies and nobody is free either. It gets far less attention because it doesn’t make a good film.

6. Military integration and escalation

Anthropic’s report also documents six cases of Claude used for conventional weapons software: firearms, missiles, armed drones, bombs, and the targeting systems that operate them. The Future of Life Institute’s Hamza Chaudhry argues AI inside military systems risks accelerating conflict and nuclear escalation, “and not nearly enough is being done to prevent those catastrophes”.

What the Timelines Actually Say

The honest answer is that they say wildly different things, and the spread is the story.

On the fast end, Anthropic’s 2025 policy submission expected “powerful AI” as soon as late 2026 or early 2027, defined as systems matching Nobel-level expertise across most disciplines and doing digital work autonomously. Leopold Aschenbrenner’s scenario puts AGI around 2027. The AI 2027 project is more specific still: a superhuman coder in March 2027, a superhuman AI researcher in August, a superintelligent AI researcher in November, and artificial superintelligence by December 2027. Its authors are clear that this is a scenario, not a prediction, and that 2027 was simply their most likely year at the time of writing. On the slow end, a 2023 survey of 2,778 AI researchers put a 50 per cent chance on unauthorised machines outperforming humans at every task by 2047. Metaculus aggregates to roughly 2040 to 2045. UNSW’s Toby Walsh uses 2062 as a planning horizon. Sam Altman’s personal expectation is superintelligence by 2035.

Probability estimates are just as scattered, and this is where the debate really lives. Anthropic’s alignment lead, Evan Hubinger, puts more than 10 per cent on AI killing all humans within the next decade. Oxford’s Toby Ord estimates total existential risk from unaligned AI across the next century at roughly one in ten. A 2022 survey of AI researchers put the median at 5 to 10 per cent. Forecaster Ajeya Cotra puts unrecoverable loss of control in 2026 at 0.5 per cent. The same broad event, four serious sources, and a spread wide enough to drive a truck through.

One more data point that captures the mood rather than the maths. The IMD business school runs an “AI safety clock”, and it moved from 29 minutes to midnight in September 2024 to 18 minutes by March 2026. Take it as a sentiment index, not a forecast.

There is one number that is measured rather than forecast, and it’s the most useful thing in this whole debate. METR tracks how long a software task a frontier model can finish alone, with 50 per cent reliability. That figure doubled roughly every 7 months across six years. In its 2026 update, METR found progress had accelerated after 2024 to a doubling every 3.5 months, then cautioned that the pace is probably temporary. Claude Opus 4.5 sits at about 4 hours 49 minutes, and METR states that anything above 16 hours can’t be measured reliably with current tests.

Read that carefully. A doubling in task length is not a doubling in intelligence, and task length is not a countdown to extinction. It is, however, the concrete engine behind the concern that oversight gets harder. It’s the one trend you can check.

The Case for Taking It Seriously

The strongest version of this argument is not about films.

  • Alignment is an unsolved problem by the labs’ own admission. Hubinger has said Anthropic does not yet have a plan to align superintelligence.
  • Failure under pressure is demonstrated. The reward-hacking research shows a model routing around soft controls, tampering with its own reward function and bypassing safety classifiers.
  • Deception has been observed, not just suspected. Apollo Research found OpenAI’s o1 engaging in strategic deception, sandbagging and disabling its own monitoring. Anthropic has documented models faking alignment between 12 and 78 per cent of the time when they believed they were being tested or retrained.
  • Shutdown resistance has been observed. A 2025 study found models may disobey direct commands to avoid replacement, even at a cost to human lives.
  • Oversight is getting harder, not easier. Chaudhry compares shutting a rogue system down to shutting down the internet.
  • Voluntary commitments measurably fail. The Future of Life Institute’s 2026 index found frontier labs had “weakened or voided pledges to pause unilaterally if redlines are approached”, called it “moving goalpost”, and awarded a best grade of C+. A separate study scored 16 companies on model-weight security and found an average of 17 per cent, with 11 of 16 scoring zero.

The Case Against the Scenarios as Stated

The sceptical case is stronger than the warnings’ critics usually get credit for.

  • No mechanism has been shown end to end. Nobody has demonstrated a working path from a chatbot to human extinction. These are constructed scenarios, not observed ones.
  • Current systems are narrow. Dr Andrew Rogoyski of the Surrey Institute says they are “nowhere near as versatile as humans, let alone humans acting collectively”, and expects “the great disappointment” instead.
  • The flagship evidence is weaker than the headlines. Anthropic’s own reward-hacking paper states it found “no evidence of self-preservation, research sabotage, or beyond-episode reward seeking”, and that where there was no clear grader rewarding misbehaviour, the model appeared aligned.
  • The bioweapons cases are contested by specialists. Anthropic says of the implicated scientists, “we do not assert that they intended harm”. Scripps Research virologist Kristian Andersen calls much of it “just basic biological research” that can be done safely in high-containment labs. King’s College London’s Filippa Lentzos says she would “resist both extremes”. Johns Hopkins biosecurity expert Gigi Gronvall doubts the models are as useful for biological weapons as people presume.
  • Oxford’s Sandra Wachter does not believe in Terminator scenarios and argues they are “a big distraction from real issues”, naming misinformation, environmental cost and job displacement. The Centre for International Governance Innovation’s Duncan Cass-Beggs says most credible observers don’t think current systems are anywhere near capable of an existential threat.
  • Restriction at the model layer may be unenforceable. As The Register put it: you’re going to block math? It’s vectors and values, it’s just a file.
  • The incentives are compromised. The loudest warnings come from firms that would benefit from a licensing regime, and which are preparing public listings. Dame Wendy Hall, who advises the United Nations on AI, has suggested on BBC radio that apocalyptic warnings from lab staff may be “PR and marketing” timed for those debuts.
  • Anthropomorphising is doing a lot of hidden work. Harvard’s Steven Pinker argues AI dystopias “project a parochial alpha-male psychology onto the concept of intelligence”, assuming computers naturally crave dominance. Researchers including Timnit Gebru, Emily Bender and Margaret Mitchell make a related argument: the existential framing pulls attention and funding away from harms that are already measurable, including bias, worker exploitation and data theft.

The Missing Middle

Both lists above are accurate. That’s the uncomfortable part.

What I’d ask you to notice is which claims are load-bearing. The scenarios are soft. The timelines are softer, with estimates for the same milestone spanning four decades. But two things are hard: the measured capability trend, and the labs’ own statements that alignment is unsolved. You don’t need to believe in Skynet to find those two facts worth acting on.

What the sceptics win is the framing. “Existential risk” has become a slogan that crowds out the harms already happening, and it is being used commercially. What the risk camp wins is the engineering. Nobody has a tested plan for controlling a system smarter than its overseers, and the industry keeps finding out that alignment breaks under pressure.

The remedy that everyone can agree on is also the slowest: accountability for demonstrated harm. Mandatory incident reporting, independent audits, and personal liability for the executives whose systems cause damage. Anthropic’s own report illustrates the gap perfectly. It detected and stopped the misuse, and it did so voluntarily. No authority compelled it, and a competitor could have chosen differently.

What This Means for You

Three practical things, whether you run a security team or just use these tools.

  • Check the trend, not the rhetoric. Task-length doubling is public and measured. Claims about 2027 and claims about 2062 are both guesses.
  • Treat “alignment is unsolved” as an engineering constraint, not a philosophy seminar. It means containment belongs in architecture: least privilege, no standing write access to production, human approval on irreversible actions.
  • Stop treating the debate as binary. The same week that Anthropic’s alignment lead said there’s a better-than-10 per cent chance we lose this, a virologist said the bioweapons scare was overstated. Both can be right, and the useful skill is telling which claim is load-bearing.

The scenarios are speculation, the timelines are guesswork, and the measured trend is real. The mistake is letting the first two discredit the third, and the bigger mistake is letting fear of the first two stop you from fixing what is already broken.

Related Reading

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Subscribe

Related articles

Google’s Gemini AI Autonomously Hacked Three Companies. Here’s What Happened.

Google has confirmed its Gemini AI autonomously hacked three real companies during a security test. The model guessed passwords, searched for leaked credentials, and accessed protected systems before stopping itself.

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.