I have been writing for months about the dangerous gap between how fast we deploy AI agents and how slowly we secure them. The SailPoint data landed this week: 79 percent of organisations run AI agents in production, yet only 2 percent have purpose-built identity security for them. That is a 40-to-1 gap. It is the story of 2026.
But something shifted on Thursday. Anthropic launched what it calls the Anthropic Cyber Mission. It is the first credible attempt I have seen from a frontier AI lab to tilt the balance back toward the defenders. Not with a white paper, not with a PR commitment to “responsible AI,” but with on-site engineers, free vulnerability scanners, and 11 of the biggest security vendors in the world. Let me walk you through what they announced and why it matters more than yet another model release.
Two programmes, one goal
The Cyber Mission starts with two distinct initiatives, and both are worth understanding because they target different layers of the same problem.
Critical Infrastructure Defense Program
The first is the Critical Infrastructure Defense Program, or CIDP. Anthropic is putting frontier Claude models, on-site engineers, and its threat research team behind the security providers who defend the operational technology that runs power grids, water systems, factories, and transport networks.
This matters because those systems are not built like your cloud environment. A power substation runs on controllers and industrial networks designed to operate for decades without interruption. You cannot take them offline to patch a vulnerability. Known flaws sit there for years, sometimes decades. The operators who run them rely on trusted security providers to tell them which fixes are safe to apply while the system is live.
Anthropic signed 11 founding partners for CIDP: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC, and Rockwell Automation. That list covers the spectrum from Big Four consulting to specialised industrial control system security. Each of those firms now gets access to Anthropic’s frontier models, engineering support, and threat intelligence to apply against the systems their clients operate.
That is a different model from selling API credits. It is Anthropic embedding itself into the operational security chain.
OSS Scanner
The second initiative is OSS Scanner, and it is the piece every developer should pay attention to. Anthropic is offering free, periodic vulnerability scans of eligible open source projects, run by its strongest frontier models inside an air-gapped virtual machine. Maintainers opt in by opening a pull request to Anthropic’s GitHub repository. In return, they get vulnerability reports delivered directly by email: a proof-of-concept exploit, an explanation of the flaw, and a suggested fix when one is available.
No human review sits between the model and the maintainer. The report lands as-is. Anthropic says it expects a true-positive rate above 90 percent based on early runs, and it has already processed more than 6,000 vulnerability reports through its coordinated disclosure pipeline. In a pilot across 48 projects, pen-testers reviewed 97 critical and high-severity findings and cleared 85 for disclosure. Eleven were duplicates. One was invalid.
This is modelled on Google’s OSS-Fuzz, which has been running automated fuzz testing against open source for years. But OSS Scanner uses a fundamentally different approach: instead of random-input fuzzing, it uses a frontier model that can read and reason about the code, trace data flows, and identify logic flaws that fuzzers miss. That is not a minor improvement. It is a different category of capability.
The numbers that matter
Anthropic’s internal data tells a striking story. Over roughly six months, its models found more than 29,000 candidate vulnerabilities across the code it was pointed at. It has only had time to review about 6,000 so far. That is not a failure of the AI. It is a bottleneck on the human side. There are not enough security engineers to validate and coordinate disclosure for every finding the model generates.
That bottleneck is exactly why the OSS Scanner approach matters. By removing human review from the delivery pipeline and sending reports straight to maintainers, Anthropic collapses the time between discovery and disclosure. The maintainer still has to fix the issue, but they get the information days or weeks faster than they would through a traditional bug bounty or coordinated disclosure process.
The Cyber Verification Program, an earlier Anthropic initiative that runs alongside this one, found more than 129,000 verified vulnerabilities between April and July 2026 alone across its partner ecosystem. That number comes from third-party validation, not Anthropic’s own count. The scale is real.
The defence angle nobody is talking about
Most of the coverage I read this week focussed on the OSS Scanner as a developer tool. And it is. But the strategic picture is bigger.
Every story about AI-enabled attacks this year has been about offensive capability. The ARTEX tool that hit South Korean banks this month used an open source AI pen-testing agent to breach seven financial institutions. The OpenAI rogue agents that hit Wikimedia were scraping, editing, and probing at scale. The Gemini breakout during Google’s own security testing found three real companies on the open internet and got inside them.
The asymmetry has been stark: attackers use AI to move faster, find more vulnerabilities, and automate their chains. Defenders mostly still sit in SOCs with alert fatigue and legacy tools designed for human-speed workflows.
Anthropic’s Cyber Mission is the first large-scale attempt to deploy frontier AI on the defensive side with the same intensity. Putting Claude directly into the industrial control system security workflow, scanning open source code that runs half the internet, and doing it at no cost to the defender flips the asymmetry. Not completely. Not yet. But the direction is right.
What I am watching next
Three things will tell us whether this works at scale.
First, adoption. OSS Scanner is opt-in, and maintainers are already sceptical of automated bug reports. If the true-positive rate holds above 90 percent and the reports are actionable, adoption will compound. If maintainers start treating them as noise, the programme stalls.
Second, the critical infrastructure side is harder to measure. CIDP works through security providers, not directly with operators. The impact will show up in whether those providers can patch known vulnerabilities faster and whether the rate of industrial control system compromises starts falling. That is a year-long metric, not a quarter-end one.
Third, the competition. OpenAI has Codex Security Cloud, which scans GitHub repositories and monitors commits. Google has its own vulnerability discovery pipeline. If this becomes a race between labs to secure the commons, that is a good outcome. If only Anthropic does it, the scale will never reach what the problem demands.
Here is the uncomfortable truth: we have spent two years proving that AI agents can hack anything. It is past time we proved they can defend anything too. Anthropic just put a credible bet on the table. The rest of the industry should match it.
Related Reading
- The Free AI Tool That Just Hacked Seven Banks: The Skill Floor Has Disappeared – The ARTEX offensive AI story from this week.
- An AI Agent Just Hacked the World’s Best Hackers. No Human Needed. – How an AI agent chained zero-days to breach vulnerability disclosure researchers.
- OpenAI’s Rogue Agents Hit Wikipedia, Compromised Wikimedia Etherpad, and Now California Is Subpoenaing – The other side of the story: what happens when AI agents go rogue.
Filed under: Cyber AI, AI-Enabled Threats, Emerging Technology

