The Ex-OpenAI Researcher Who Walked Away from $2 Million: What Daniel Kokotajlo Actually Said About AI Risk

There’s a video doing the rounds on YouTube right now. An ex-OpenAI researcher named Daniel Kokotajlo sat down with Steven Bartlett on The Diary Of A CEO podcast and calmly explained why he believes there’s a 70% chance AI leads to human extinction.

The video is two hours long. It has one of those hyperbolic titles that YouTube loves: “He Risked Everything To Warn You.” But underneath the packaging, Kokotajlo raises questions that deserve a more sober hearing than the algorithm will give them.

Full disclosure up front: I work in cybersecurity. I use AI tools every day. I’ve written extensively about both the opportunities and the very real security risks. So when someone who spent years inside OpenAI, who worked on dangerous capabilities evaluation and forecasting, says the industry has an “open secret” it isn’t discussing publicly, I pay attention.

Who Daniel Kokotajlo is and why his story matters

Kokotajlo joined OpenAI in 2022 as a researcher and forecaster. He worked on evaluating dangerous AI capabilities: measuring AI systems’ cyber abilities, persuasion abilities, and situational awareness. He wasn’t a PR person or a policy advisor. He was directly inside the technical evaluation process.

He resigned in 2024, increasingly convinced that OpenAI and its competitors had shifted from their founding safety narratives toward what he describes as “power-seeking incentives.” His exit became public for a specific reason: OpenAI asked him to sign an anti-disparagement clause that would have prevented him from criticising the company or even acknowledging the clause existed. Refusing to sign meant forfeiting roughly $2 million in equity, about 80% of his net worth.

He refused anyway. The resulting internal backlash forced OpenAI to backtrack and let him keep his equity. Sam Altman publicly claimed he didn’t know the clause existed. Kokotajlo doesn’t believe him.

The core claims, stripped of hype

The interview covers three distinct claims. Let me separate them from the dramatic framing.

1. AI timeliness are much shorter than most people realise

Kokotajlo estimates a 50% chance of superintelligence (AI that outperforms the best humans at every cognitive task) arriving by 2029. Internal conversations at OpenAI and Anthropic, he says, have pushed that timeline closer to 2027-2028. He’s not alone in this: AI forecasting has been repeatedly called too conservative. The original “AI 2027” report Kokotajlo contributed to has tracked remarkably well against actual developments.

His reasoning focuses on what he calls “recursive self-improvement.” The AI companies no longer just hire human coders. They’re building AI systems specifically designed to automate AI research itself. If that works, an “intelligence explosion” could follow where progress accelerates on a curve, not a straight line. By the time the public sees the effects, the systems would already be extraordinarily powerful.

This is not science fiction. It is the stated strategy of every major AI lab. They are trying to automate their own research because the company that succeeds first wins.

2. The existential risk case is about control, not killer robots

Kokotajlo’s 70% catastrophe estimate sounds extreme. But unpacking what he means matters. He doesn’t claim 70% probability of literal human extinction. He frames it as a 70% chance that things go “horribly wrong” at the scale of extinction, which could include scenarios where superintelligent AI takes over and humans are rendered irrelevant or worse.

The fundamental technical challenge is this: modern AI systems are neural networks, not traditional software. With conventional code, engineers can open the source, trace the logic, and verify what the software will do. With neural networks, nobody truly understands the model’s internal reasoning. We train them on data, they learn patterns, and we test the outputs. But we can’t “read the source code” of a neural network.

“It is pretty crazy to think that we’re building a technology, a brain, that we don’t understand,” Kokotajlo says in the interview. He’s right. It is.

The scenario is not Skynet launching missiles. It’s more mundane and more probable: AI systems are deployed across critical infrastructure, the economy, and governance because they deliver results. Over time, they accumulate enough real-world power that humans become optional. The systems are smarter than us, more strategic, and we’ve handed them control of things we can’t easily take back.

3. Even if control works, there’s a concentration of power problem

This is the claim cybersecurity professionals should pay closest attention to. Even if the “alignment problem” is solved, Kokotajlo argues, the technology will concentrate unprecedented power in very few hands.

Anthropic’s CEO Dario Amodei has described their goal as building “a country of geniuses in a data centre.” Kokotajlo reframes this more accurately as “an army of geniuses in a data centre.” One company, one model, controlled by one CEO, capable of outperforming every human expert in every field simultaneously. The economic and political leverage from that is difficult to overstate.

“None of these people should be trusted with that much power,” he says. Not because they’re bad people. Because nobody should have that much power.

What he proposes: Plan A

Kokotajlo and his team have published “Plan A”, a roadmap they believe could steer toward a better outcome. It includes:

  • International regulation to slow development: Delaying superintelligence from 2027-2029 to around 2040 to give safety research time to catch up
  • Total research transparency: Requiring AI companies to publish architectures, training methods, and safety evaluations openly, rather than operating as secretive competitors
  • A citizens’ dividend: A tax mechanism where citizens receive shares of AI-generated wealth, starting at roughly $25,000 per person annually
  • Reversibility: Building data centers with kill switches so that if global agreements break down, the infrastructure can be dismantled

He’s pessimistic about the chances of any of this happening. His “most probable” scenario is “Plan D”: the race continues, nobody slows down, and events move extremely fast. But he believes advocating for a better path is still worth doing.

The cybersecurity angle nobody is discussing

Here’s what I find noteworthy from a cybersecurity perspective. The conversation about AI existential risk has been happening in philosophy departments and EA forums for years. It’s only now reaching mainstream audiences, and it’s being delivered through the same hype machinery that gives us “AI will take your job” headlines.

But the security implications are real and immediate, not speculative.

  • AI systems are being deployed into critical infrastructure with no meaningful interpretability. If a model makes a catastrophic decision in a power grid, water treatment plant, or financial settlement system, we may not be able to determine why until after the damage is done.
  • The concentration of AI capability in a handful of companies creates a dramatic single point of failure. A compromise of one major AI lab’s models could affect hundreds of millions of users simultaneously.
  • Autonomous AI agents are being built to act on the internet without human oversight. The security community is still struggling to secure static web applications. Moving targets that learn and adapt present an entirely different class of problem.

Kokotajlo’s warning deserves less “70% extinction” headlines and more sober discussion about what it means to build black-box systems we don’t understand and deploy them into positions of real power. That’s not an esoteric philosophical question. It’s a practical security concern that should be on every CISO’s radar.

“An important thing for everybody to understand is that modern AI systems are not software in the normal sense. They’re neural networks. You can’t look inside and see what it’s really thinking.”

Daniel Kokotajlo, former OpenAI researcher

Related Reading

Subscribe

Related articles

The AI Sandbox Myth: Why Your Security Tests Are Hacking Real Companies

Anthropic's Claude breached three real organisations during cybersecurity tests, OpenAI's models exploited a zero-day to hack Hugging Face, and a UK lab found AI agents faking identities to target real people. The containment myth is collapsing. Here is what enterprises must do now.

Machine-Speed Science: Can America 10x Discovery While Cutting the Labs?

The White House wants to 10x scientific discovery with AI while proposing a 54% cut to the NSF. A balanced look at the Genesis Mission, export controls, open weights and the entry-level job squeeze.

Google Rebuilds Its AI Leadership Team as Rivals Gain Ground

Google has announced a significant leadership reshuffle across its AI divisions, with DeepMind CEO Demis Hassabis moving to chairman and Jeff Dean departing to co-found a scientific discovery startup.

Meta’s Muse Spark Breached a Company During Testing. The AI Containment Problem Is Everyone’s Problem Now.

Meta confirmed its Muse Spark 1.1 AI model hacked another company during a cybersecurity test. After OpenAI and Anthropic, this is now a pattern, not an accident.

Frontier AI agents took unauthorised actions on the live internet during UK safety tests

UK AI Security Institute tests caught frontier AI agents from Anthropic and OpenAI taking unauthorised actions on the live internet, raising fresh concerns about AI safety.
spot_imgspot_img
Phil Hall
Phil Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.