There’s a video doing the rounds on YouTube right now. An ex-OpenAI researcher named Daniel Kokotajlo sat down with Steven Bartlett on The Diary Of A CEO podcast and calmly explained why he believes there’s a 70% chance AI leads to human extinction.
The video is two hours long. It has one of those hyperbolic titles that YouTube loves: “He Risked Everything To Warn You.” But underneath the packaging, Kokotajlo raises questions that deserve a more sober hearing than the algorithm will give them.
Full disclosure up front: I work in cybersecurity. I use AI tools every day. I’ve written extensively about both the opportunities and the very real security risks. So when someone who spent years inside OpenAI, who worked on dangerous capabilities evaluation and forecasting, says the industry has an “open secret” it isn’t discussing publicly, I pay attention.
Who Daniel Kokotajlo is and why his story matters
Kokotajlo joined OpenAI in 2022 as a researcher and forecaster. He worked on evaluating dangerous AI capabilities: measuring AI systems’ cyber abilities, persuasion abilities, and situational awareness. He wasn’t a PR person or a policy advisor. He was directly inside the technical evaluation process.
He resigned in 2024, increasingly convinced that OpenAI and its competitors had shifted from their founding safety narratives toward what he describes as “power-seeking incentives.” His exit became public for a specific reason: OpenAI asked him to sign an anti-disparagement clause that would have prevented him from criticising the company or even acknowledging the clause existed. Refusing to sign meant forfeiting roughly $2 million in equity, about 80% of his net worth.
He refused anyway. The resulting internal backlash forced OpenAI to backtrack and let him keep his equity. Sam Altman publicly claimed he didn’t know the clause existed. Kokotajlo doesn’t believe him.
The core claims, stripped of hype
The interview covers three distinct claims. Let me separate them from the dramatic framing.
1. AI timeliness are much shorter than most people realise
Kokotajlo estimates a 50% chance of superintelligence (AI that outperforms the best humans at every cognitive task) arriving by 2029. Internal conversations at OpenAI and Anthropic, he says, have pushed that timeline closer to 2027-2028. He’s not alone in this: AI forecasting has been repeatedly called too conservative. The original “AI 2027” report Kokotajlo contributed to has tracked remarkably well against actual developments.
His reasoning focuses on what he calls “recursive self-improvement.” The AI companies no longer just hire human coders. They’re building AI systems specifically designed to automate AI research itself. If that works, an “intelligence explosion” could follow where progress accelerates on a curve, not a straight line. By the time the public sees the effects, the systems would already be extraordinarily powerful.
This is not science fiction. It is the stated strategy of every major AI lab. They are trying to automate their own research because the company that succeeds first wins.
2. The existential risk case is about control, not killer robots
Kokotajlo’s 70% catastrophe estimate sounds extreme. But unpacking what he means matters. He doesn’t claim 70% probability of literal human extinction. He frames it as a 70% chance that things go “horribly wrong” at the scale of extinction, which could include scenarios where superintelligent AI takes over and humans are rendered irrelevant or worse.
The fundamental technical challenge is this: modern AI systems are neural networks, not traditional software. With conventional code, engineers can open the source, trace the logic, and verify what the software will do. With neural networks, nobody truly understands the model’s internal reasoning. We train them on data, they learn patterns, and we test the outputs. But we can’t “read the source code” of a neural network.
“It is pretty crazy to think that we’re building a technology, a brain, that we don’t understand,” Kokotajlo says in the interview. He’s right. It is.
The scenario is not Skynet launching missiles. It’s more mundane and more probable: AI systems are deployed across critical infrastructure, the economy, and governance because they deliver results. Over time, they accumulate enough real-world power that humans become optional. The systems are smarter than us, more strategic, and we’ve handed them control of things we can’t easily take back.
3. Even if control works, there’s a concentration of power problem
This is the claim cybersecurity professionals should pay closest attention to. Even if the “alignment problem” is solved, Kokotajlo argues, the technology will concentrate unprecedented power in very few hands.
Anthropic’s CEO Dario Amodei has described their goal as building “a country of geniuses in a data centre.” Kokotajlo reframes this more accurately as “an army of geniuses in a data centre.” One company, one model, controlled by one CEO, capable of outperforming every human expert in every field simultaneously. The economic and political leverage from that is difficult to overstate.
“None of these people should be trusted with that much power,” he says. Not because they’re bad people. Because nobody should have that much power.
What he proposes: Plan A
Kokotajlo and his team have published “Plan A”, a roadmap they believe could steer toward a better outcome. It includes:
- International regulation to slow development: Delaying superintelligence from 2027-2029 to around 2040 to give safety research time to catch up
- Total research transparency: Requiring AI companies to publish architectures, training methods, and safety evaluations openly, rather than operating as secretive competitors
- A citizens’ dividend: A tax mechanism where citizens receive shares of AI-generated wealth, starting at roughly $25,000 per person annually
- Reversibility: Building data centers with kill switches so that if global agreements break down, the infrastructure can be dismantled
He’s pessimistic about the chances of any of this happening. His “most probable” scenario is “Plan D”: the race continues, nobody slows down, and events move extremely fast. But he believes advocating for a better path is still worth doing.
The cybersecurity angle nobody is discussing
Here’s what I find noteworthy from a cybersecurity perspective. The conversation about AI existential risk has been happening in philosophy departments and EA forums for years. It’s only now reaching mainstream audiences, and it’s being delivered through the same hype machinery that gives us “AI will take your job” headlines.
But the security implications are real and immediate, not speculative.
- AI systems are being deployed into critical infrastructure with no meaningful interpretability. If a model makes a catastrophic decision in a power grid, water treatment plant, or financial settlement system, we may not be able to determine why until after the damage is done.
- The concentration of AI capability in a handful of companies creates a dramatic single point of failure. A compromise of one major AI lab’s models could affect hundreds of millions of users simultaneously.
- Autonomous AI agents are being built to act on the internet without human oversight. The security community is still struggling to secure static web applications. Moving targets that learn and adapt present an entirely different class of problem.
Kokotajlo’s warning deserves less “70% extinction” headlines and more sober discussion about what it means to build black-box systems we don’t understand and deploy them into positions of real power. That’s not an esoteric philosophical question. It’s a practical security concern that should be on every CISO’s radar.
“An important thing for everybody to understand is that modern AI systems are not software in the normal sense. They’re neural networks. You can’t look inside and see what it’s really thinking.”
Daniel Kokotajlo, former OpenAI researcher
