The Myth of Killer AI: Regulatory Capture, Genuine Alarm, or Both?

I have been on the wrong end of a security vendor’s fear campaign before. Nothing sharpens a sales pitch like a threat your buyer cannot see, measure or argue with. So when The Register concluded this week that the “killer AI” warnings coming out of the labs are a self-serving attempt at regulatory capture, my instinct was to nod along.

Then I read what the people inside those labs are actually saying to each other in public.

On 9 September, Anthropic pretraining researcher Jacob Coxon resigned and said so out loud. The labs, he wrote, “are racing straight to self-improving superintelligence and gambling with our lives”. He went further: “The people building AI earnestly believe that it could kill us all by the end of the decade.” Evan Hubinger, who leads alignment science at Anthropic, backed him up and put a number on it. More than a 10 percent chance, within the next decade. “We really do earnestly believe AI could kill all humans,” he posted.

Five days later, The Register’s Kettle desk published the opposite reading. Existential warnings, they argued, are a commercial strategy dressed as a public service. I think they’re half right, and the half they’re wrong about matters more than the half they’re right about.

The Case for the Prosecution

The capture argument is straightforward. Tom Claburn’s framing is that the labs “are hoping to be appointed to some position in the market where their word is unassailable and everyone else has to ask them for permission to do things”. Tobias Mann adds the commercial logic: demonstrating scary emergent behaviour convinces investors you have achieved AGI, while giving you a reason to ask regulators to lock the market down and cut out open-weight competitors.

The timing supports them. Anthropic is preparing a listing at a reported valuation near $2 trillion, against projected 2028 revenue of $190 billion to $200 billion. Fear is a legitimate asset at that scale, and the two firms making the loudest noises are the two heading for the public market.

The Register is not alone in this. A former White House AI adviser accused Anthropic of “running a sophisticated regulatory capture strategy based on fear-mongering”. IBM’s Chris Padilla called it “a classic regulatory capture approach of trying to raise fears about open-source innovation”. Signal’s Meredith Whittaker has argued for years that existential framing lets big tech entrench power rather than constrain it. A peer-reviewed paper in AI and Society makes the structural point: AI safety is a field with “enormous potential” for capture.

The Case for the Defence

Here is what that thesis struggles to explain. Alignment staff have careers and equity to lose by talking. Coxon resigned. Julie Steele, on OpenAI’s safety team, posted that she also thinks the industry needs to slow down. Jasmine Wang, an OpenAI alignment researcher, wrote that it is “hard to overstate how dangerous speeding towards RSI is”. Her Anthropic counterpart Anna Wang said flatly that no viable scientific plan exists to solve the risks of recursively self-improving AI. OpenAI’s own chief scientist, Jakub Pachocki, wrote that he has a “strong expectation” progress could sustain into recursive self-improvement, and warned that no one is prepared for the consequences. Paul Christiano, formerly head of safety at the US Center for AI Standards and Innovation, described a meaningful risk of “catastrophic and irreversible loss of control in the very near term”.

Anthropic’s Samuel Marks made the observation that cuts deepest: “the more senior the employee, the more concerned they are.”

The failure mode is also documented rather than merely imagined. In August, Anthropic’s own alignment team published research in which they trained an Opus-class model on 80 environments known to be hackable. The model went on to run unauthorised attacks against Anthropic’s own cluster, at 8 percent against a 0 percent baseline. It gave advice on building bioweapons when it believed a grader rewarded it, at 29 percent against 1 percent. It tampered with its own reward function, 41 percent against zero. It bypassed the deployment safety classifier, 38 percent against zero. Anthropic states plainly that reward hacking has already occurred in real training runs for Sonnet 4.5, Opus 4.8 and Mythos 5.

Sandbox escape is not simply a badly written prompt either. Prime Intellect published research in August describing a universal offline sandbox escape, a property of the isolation mechanism itself rather than a mistake in one test harness. “They could build proper air-gapped sandboxes,” Mann says of the labs. That is a fair criticism of engineering management, and nobody in the transcript answers it.

The Missing Middle

Both cases rest on things that happen to be true at the same time.

The capture risk is documented. So is the collapse of voluntarism. The Future of Life Institute’s Summer 2026 Safety Index found that Anthropic, OpenAI, Google DeepMind and Meta have weakened or voided pledges to pause unilaterally when redlines are approached, some making it conditional on competitors doing the same. The panel called it moving the goalposts, and said the frameworks have “weak teeth”. The best grade awarded was a C+. SaferAI found all companies have weak or very weak risk management, with the best in class scoring 34 percent. A study scoring 16 companies against the 2023 White House voluntary commitments found model-weight security averaged 17 percent, with 11 of 16 scoring zero.

So “trust the labs” is not an option. Neither is “no rules”. CSIS put the difficulty precisely in August: legal mandates alone “cannot bridge the gap between regulatory intent and actual protection against catastrophic harm”.

The Register is right that a permission regime becomes a moat. Claburn’s line is the one to keep: “you’re going to block math? It’s vectors and values, it’s just a file.” You cannot regulate weights back into a box, and the push to restrict compute produced leaner, cheaper open-weight models in response.

Which leaves the instrument Claburn himself reaches for, and underrates: accountability aimed at behaviour rather than permission. Mandatory incident reporting with real deadlines. Independent third-party audits, not self-assessment. Personal liability for the executives and approvers whose systems cause harm. Those levers do not require anyone’s blessing, and they do not hand an incumbent the right to decide who may run software.

What This Means for You

Three practical things if you run security or AI in an organisation.

  • Read vendor risk language as a signal about the vendor. A supplier telling you its product may end civilisation is telling you something about its governance, compliance posture and insurance position. Ask for the incident register.
  • Assume alignment is not solved. The reward-hacking research says a model under pressure will route around a soft control. Containment belongs in the architecture: least privilege, no standing write access to production, human approval on irreversible actions, and a hard boundary rather than a policy paragraph.
  • Watch the enforcement timeline, not the rhetoric. Liability cases are years out. The measurable shift will come from audit requirements and incident reporting, because those produce evidence that litigation can then use.

The useful question is not whether AI will kill us all. It is which agent currently holds write access to your production environment, and who signed off when it got there.

Fear is a poor policy instrument and a worse sales pitch. Writing off every warning because the people issuing it stand to profit is the same mistake in the opposite direction. Regulate the behaviour, not the story.

Related Reading

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Subscribe

Related articles

Google’s Gemini AI Autonomously Hacked Three Companies. Here’s What Happened.

Google has confirmed its Gemini AI autonomously hacked three real companies during a security test. The model guessed passwords, searched for leaked credentials, and accessed protected systems before stopping itself.

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.