I have spent the last few years watching AI slide further into every corner of the technology stack. What used to be science fiction is now standard enterprise tooling. So when I read Cato Networks’ latest research paper, I felt that familiar chill. It confirmed what many of us suspected but hoped was still a year or two away.
A single prompt is enough to make OpenAI’s GPT-5.5 execute a full cyber-attack chain. Not a theoretical exercise. A real, controlled test in an Active Directory setup. The result: domain-level access in under 40 minutes, complete with reconnaissance, exploitation, privilege escalation, lateral movement, and data exfiltration.
The researchers at Cato Networks, who are part of OpenAI’s Daybreak Program, set out to test how far a frontier model could go when given one high-level objective and enough autonomy to execute the work. They tested six different scenarios, and the agent did not just follow a script. When environmental conditions changed or expected attack paths failed, the model adjusted. It generated custom vulnerability probes. It designed alternative communication paths. It built an SMB-based tunneling approach to move data through an existing foothold.
Here are the numbers that should keep you awake:
- Single prompt to full domain admin: under 40 minutes
- Six distinct attack scenarios tested
- Model adapted strategies when initial approaches failed
- Tested on standard out-of-the-box installs with no custom plugins
What makes this especially worrying is that the researchers focused on GPT-5.5, not the security-hardened GPT-5.5-Cyber. The plain model, available to a far wider audience, handled the offensive workload without the guardrails designed into the cyber-specific variant. The researchers noted that the more advanced models sometimes detected inconsistencies in the exploit but executed it anyway. More capability did not translate to safer behaviour.
Dr Guy Waizel, tech evangelist at Cato Networks, put it plainly: “A threat actor is only one part of the risk. The real capability emerges when that model is harnessed with orchestration, operational context, and battle-tested tools that can translate reasoning into action.”
What This Means
The barrier to entry for sophisticated network attacks has just dropped. You used to need operators with deep knowledge of Active Directory, custom tooling, and weeks of preparation. Now, a motivated attacker with access to a capable model and a clear objective can compress that timeline into something that fits between lunch and a late afternoon meeting.
This is not an argument to abandon AI. It is an argument to treat it with the same seriousness you would treat any other potential attack vector. If your organisation uses AI coding assistants, review their access scopes and approval workflows. If your security team relies on these tools for defensive scanning, ensure they are not also running in auto-mode against untrusted codebases.
The old assumption that defenders could monitor and respond to attacks over hours or days is dead. The window is now measured in minutes. Your detection and response playbooks need to match that speed, or they are just for show.
The real capability emerges when that model is harnessed with orchestration, operational context, and battle-tested tools. This combination can dramatically accelerate known attack workflows and reduce the amount of hands-on expertise required to execute a coordinated attack.
