OpenAI Fired Safety Researchers Hit Back: Culture Is ‘Chilling’

Three weeks after OpenAI publicly pledged to open its doors to outside safety experts, the company fired three of its own safety researchers. Now those researchers are hitting back, and their warnings deserve attention.

On October 9, researchers Tomek Korbak, Jasmine Wang, and Mikita Balesni published an open letter denying the company’s claims that they mishandled sensitive information. More importantly, they warned that their dismissals risk “chilling” OpenAI’s safety culture at a time when external oversight has never been more critical.

The Researchers’ Side of the Story

The trio say they acted within their job mandates the entire time. In their open letter, they argue that if past conduct is now grounds for firing, every staff member is left guessing where the line actually sits. That ambiguity, they warn, is precisely the kind of environment that discourages safety researchers from raising hard questions.

Each researcher detailed their individual case. They denied leaking a story about less monitorable AI models, insisting their outside work stayed within policy with leadership kept in the loop at every stage. Wang said her access to an executive’s email was delegated specifically for recruiting purposes, and that she reported accidentally opening a sensitive message within minutes of realising what had happened.

OpenAI’s position is different. The company told TechCrunch that the firings were not about raising safety concerns, and that the researchers had engaged in a “pattern of misconduct.” OpenAI also said it agreed with the recommendations the researchers made in their letter, but stood by its decision to dismiss them.

Why the Firings Matter for AI Safety

The substance of the researchers’ warning goes beyond their own employment. They argue that firing safety staff for conduct that was previously considered acceptable gives OpenAI cover to skip embedding independent auditors. The company has promised external oversight before, but if internal researchers can be dismissed when their work becomes inconvenient, what real check exists on the company’s safety practices?

This concern lands in a well-worn groove. OpenAI’s Superalignment team dissolved earlier this year amid internal disagreements over safety priorities. More recently, a senior safety lead departed publicly citing a “broken” culture. The pattern is becoming difficult to ignore.

Meanwhile, Anthropic has already appointed its first independent evaluator and published a framework for how external safety testing will work. OpenAI has talked about doing the same, but the rhetoric has not yet translated into concrete, verifiable action. These firings do not help that impression.

What Happens Next

The researchers are urging OpenAI to protect its open culture and ensure models remain monitorable by external parties. They want clear, written guarantees that safety staff can raise concerns without fear of retaliation.

OpenAI has not commented publicly beyond its statement to TechCrunch. The company has time to address these concerns, but the clock is ticking. If the perception takes hold that internal safety researchers are penalised for doing their jobs, attracting and retaining the talent needed to build safe frontier models becomes significantly harder.

The Bigger Picture

This story is not really about three people losing their jobs. It is about whether the structural safeguards inside the world’s most prominent AI company are working, or whether they are being quietly dismantled. The researchers’ public letter is a signal that the people closest to the safety work do not believe the system is functioning as advertised.

For anyone watching AI governance from the outside, this is a moment to pay attention. Safety culture is fragile. It depends on people feeling empowered to identify problems without calculating whether doing so will end their careers. When that calculation changes, the culture shifts, often invisibly, until a crisis reveals what was lost.

Anthropic has its evaluator. Google DeepMind has its governance structures. OpenAI has promises, three fired researchers, and growing questions about whether the gap between rhetoric and reality is widening rather than closing.


Want to go deeper? Read the original open letter from Korbak, Wang, and Balesni, or follow TechCrunch’s coverage of OpenAI’s response. For more on AI safety culture, revisit our coverage of the Superalignment team’s dissolution and the senior safety lead’s departure over a “broken” culture.

Subscribe

Related articles

Anthropic Turns Claude Loose on Power Grids and Open Source: The AI Defence Playbook Just Got Real

Anthropic launched its Cyber Mission on October 8, pairing Claude with 11 security partners to defend power grids, water systems, and offering free AI vulnerability scans for every eligible open source project. This is what it means for enterprise defenders.

Zuckerberg and Chan’s Biohub Pours $1.8 Billion Into AI That Simulates Human Cells

Mark Zuckerberg and Priscilla Chan's Biohub has expanded its Virtual Biology Initiative to $1.8 billion, backed by the US government, Google DeepMind, and Meta. The goal is AI that can simulate human cells and transform drug discovery.

The Free AI Tool That Just Hacked Seven Banks: The Skill Floor Has Disappeared

An open-source AI penetration testing tool called ARTEX was used to breach seven South Korean financial institutions and expose 68,000 customer records. The scary part is anyone can use it.

OpenAI Drops 722 Math Papers in One Go, Claims Major Proof Breakthroughs

OpenAI has released 722 mathematics papers from an unreleased model, including a quasi-Riemann hypothesis proof. The drop marks a turning point for AI-driven discovery.

Someone Built a Fake AI Ad Empire to Steal Your Login. And It Worked.

A human-operated phishing platform is impersonating ChatGPT, Gemini, Claude, Perplexity and Meta Muse with fake advertising portals that steal credentials and bypass MFA. Island researchers found hundreds of victims and the campaign is still running.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.

This site uses Akismet to reduce spam. Learn how your comment data is processed.