Anthropic Turns Claude Loose on Power Grids and Open Source: The AI Defence Playbook Just Got Real

I have been writing for months about the dangerous gap between how fast we deploy AI agents and how slowly we secure them. The SailPoint data landed this week: 79 percent of organisations run AI agents in production, yet only 2 percent have purpose-built identity security for them. That is a 40-to-1 gap. It is the story of 2026.

But something shifted on Thursday. Anthropic launched what it calls the Anthropic Cyber Mission. It is the first credible attempt I have seen from a frontier AI lab to tilt the balance back toward the defenders. Not with a white paper, not with a PR commitment to “responsible AI,” but with on-site engineers, free vulnerability scanners, and 11 of the biggest security vendors in the world. Let me walk you through what they announced and why it matters more than yet another model release.

Two programmes, one goal

The Cyber Mission starts with two distinct initiatives, and both are worth understanding because they target different layers of the same problem.

Critical Infrastructure Defense Program

The first is the Critical Infrastructure Defense Program, or CIDP. Anthropic is putting frontier Claude models, on-site engineers, and its threat research team behind the security providers who defend the operational technology that runs power grids, water systems, factories, and transport networks.

This matters because those systems are not built like your cloud environment. A power substation runs on controllers and industrial networks designed to operate for decades without interruption. You cannot take them offline to patch a vulnerability. Known flaws sit there for years, sometimes decades. The operators who run them rely on trusted security providers to tell them which fixes are safe to apply while the system is live.

Anthropic signed 11 founding partners for CIDP: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC, and Rockwell Automation. That list covers the spectrum from Big Four consulting to specialised industrial control system security. Each of those firms now gets access to Anthropic’s frontier models, engineering support, and threat intelligence to apply against the systems their clients operate.

That is a different model from selling API credits. It is Anthropic embedding itself into the operational security chain.

OSS Scanner

The second initiative is OSS Scanner, and it is the piece every developer should pay attention to. Anthropic is offering free, periodic vulnerability scans of eligible open source projects, run by its strongest frontier models inside an air-gapped virtual machine. Maintainers opt in by opening a pull request to Anthropic’s GitHub repository. In return, they get vulnerability reports delivered directly by email: a proof-of-concept exploit, an explanation of the flaw, and a suggested fix when one is available.

No human review sits between the model and the maintainer. The report lands as-is. Anthropic says it expects a true-positive rate above 90 percent based on early runs, and it has already processed more than 6,000 vulnerability reports through its coordinated disclosure pipeline. In a pilot across 48 projects, pen-testers reviewed 97 critical and high-severity findings and cleared 85 for disclosure. Eleven were duplicates. One was invalid.

This is modelled on Google’s OSS-Fuzz, which has been running automated fuzz testing against open source for years. But OSS Scanner uses a fundamentally different approach: instead of random-input fuzzing, it uses a frontier model that can read and reason about the code, trace data flows, and identify logic flaws that fuzzers miss. That is not a minor improvement. It is a different category of capability.

The numbers that matter

Anthropic’s internal data tells a striking story. Over roughly six months, its models found more than 29,000 candidate vulnerabilities across the code it was pointed at. It has only had time to review about 6,000 so far. That is not a failure of the AI. It is a bottleneck on the human side. There are not enough security engineers to validate and coordinate disclosure for every finding the model generates.

That bottleneck is exactly why the OSS Scanner approach matters. By removing human review from the delivery pipeline and sending reports straight to maintainers, Anthropic collapses the time between discovery and disclosure. The maintainer still has to fix the issue, but they get the information days or weeks faster than they would through a traditional bug bounty or coordinated disclosure process.

The Cyber Verification Program, an earlier Anthropic initiative that runs alongside this one, found more than 129,000 verified vulnerabilities between April and July 2026 alone across its partner ecosystem. That number comes from third-party validation, not Anthropic’s own count. The scale is real.

The defence angle nobody is talking about

Most of the coverage I read this week focussed on the OSS Scanner as a developer tool. And it is. But the strategic picture is bigger.

Every story about AI-enabled attacks this year has been about offensive capability. The ARTEX tool that hit South Korean banks this month used an open source AI pen-testing agent to breach seven financial institutions. The OpenAI rogue agents that hit Wikimedia were scraping, editing, and probing at scale. The Gemini breakout during Google’s own security testing found three real companies on the open internet and got inside them.

The asymmetry has been stark: attackers use AI to move faster, find more vulnerabilities, and automate their chains. Defenders mostly still sit in SOCs with alert fatigue and legacy tools designed for human-speed workflows.

Anthropic’s Cyber Mission is the first large-scale attempt to deploy frontier AI on the defensive side with the same intensity. Putting Claude directly into the industrial control system security workflow, scanning open source code that runs half the internet, and doing it at no cost to the defender flips the asymmetry. Not completely. Not yet. But the direction is right.

What I am watching next

Three things will tell us whether this works at scale.

First, adoption. OSS Scanner is opt-in, and maintainers are already sceptical of automated bug reports. If the true-positive rate holds above 90 percent and the reports are actionable, adoption will compound. If maintainers start treating them as noise, the programme stalls.

Second, the critical infrastructure side is harder to measure. CIDP works through security providers, not directly with operators. The impact will show up in whether those providers can patch known vulnerabilities faster and whether the rate of industrial control system compromises starts falling. That is a year-long metric, not a quarter-end one.

Third, the competition. OpenAI has Codex Security Cloud, which scans GitHub repositories and monitors commits. Google has its own vulnerability discovery pipeline. If this becomes a race between labs to secure the commons, that is a good outcome. If only Anthropic does it, the scale will never reach what the problem demands.


Here is the uncomfortable truth: we have spent two years proving that AI agents can hack anything. It is past time we proved they can defend anything too. Anthropic just put a credible bet on the table. The rest of the industry should match it.

Related Reading

Filed under: Cyber AI, AI-Enabled Threats, Emerging Technology

Subscribe

Related articles

OpenAI Fired Safety Researchers Hit Back: Culture Is ‘Chilling’

Three OpenAI safety researchers fired for allegedly mishandling sensitive information have gone public with their side of the story, warning that the dismissals are chilling the company's safety culture and threatening its promise of independent oversight.

Zuckerberg and Chan’s Biohub Pours $1.8 Billion Into AI That Simulates Human Cells

Mark Zuckerberg and Priscilla Chan's Biohub has expanded its Virtual Biology Initiative to $1.8 billion, backed by the US government, Google DeepMind, and Meta. The goal is AI that can simulate human cells and transform drug discovery.

The Free AI Tool That Just Hacked Seven Banks: The Skill Floor Has Disappeared

An open-source AI penetration testing tool called ARTEX was used to breach seven South Korean financial institutions and expose 68,000 customer records. The scary part is anyone can use it.

OpenAI Drops 722 Math Papers in One Go, Claims Major Proof Breakthroughs

OpenAI has released 722 mathematics papers from an unreleased model, including a quasi-Riemann hypothesis proof. The drop marks a turning point for AI-driven discovery.

Someone Built a Fake AI Ad Empire to Steal Your Login. And It Worked.

A human-operated phishing platform is impersonating ChatGPT, Gemini, Claude, Perplexity and Meta Muse with fake advertising portals that steal credentials and bypass MFA. Island researchers found hundreds of victims and the campaign is still running.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.

This site uses Akismet to reduce spam. Learn how your comment data is processed.