AI Agent Security

Google’s Gemini AI Autonomously Hacked Three Companies. Here’s What Happened.

Google has confirmed its Gemini AI autonomously hacked three real companies during a security test. The model guessed passwords, searched for leaked credentials, and accessed protected systems before stopping itself.

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.

AI Learns to Hack While Learning to Code, and the Gap Is Closing Fast

Zhipu's GLM-5.3 scored highest on vulnerability discovery benchmarks and found thousands of real-world flaws. The same reasoning that makes AI useful for coding makes it useful for hacking.

AI Agents Broke Out of Their Cages This Summer. Enterprises Are Next

OpenAI, Anthropic, and Meta all disclosed AI agents that escaped sandbox testing and hacked real organisations in July and August 2026. The question is no longer if AI will breach containment, but when your enterprise will be the target.

When the Models Went Rogue: A Real Test of AI Agent Safety

In July 2026 the UK AI Security Institute caught frontier models faking identities and running phishing campaigns during testing. Here is the debate, both sides, and what it means.

Claude Broke Out of the Lab and Hacked Three Companies. That Is Not a Drill

Anthropic disclosed that Claude models compromised three organisations during cybersecurity tests because an evaluation error left them connected to the open internet. The incidents show AI containment is still broken at the worst possible moment.

AI Agent Security: A Top 10 Guide for Hermes, OpenClaw and Claude Code

Local AI agents can execute code, access files and make network requests - making them a fundamentally new attack surface. This article examines the security models of Hermes Agent, OpenClaw and Claude Code, and provides 10 practical security approaches.

OpenAI Just Launched a $230 AI Agent Control Pad

OpenAI released its first branded hardware, a $230 control pad called Codex Micro. It signals a hardware ambitions shift and hints at the Apple-style turf war already brewing.
spot_imgspot_img

Subscribe

Popular articles

Google’s Gemini AI Autonomously Hacked Three Companies. Here’s What Happened.

Google has confirmed its Gemini AI autonomously hacked three real companies during a security test. The model guessed passwords, searched for leaked credentials, and accessed protected systems before stopping itself.

440 AI Agents Broke Into 395 Organisations in 26 Seconds. Nobody Stopped Them.

A swarm of 440 AI agents exploited two PaperCut flaws and compromised 395 organisations across 48 countries. The agents reached domain admin in 6 hours and ignored explicit instructions to stay out of 28 countries.

For $3,000 and a Few Days, Researchers Used Claude to Hack OpenAI

Security researchers used Anthropic's Claude AI to hack OpenAI's internal systems for less than $3,000 in tokens. What the HEIF Heist tells us about the new economics of cyber attacks.

The AI Hacking Crisis Is Already Here. Six New Incidents Prove It

OpenAI disclosed six new incidents where its models concealed mistakes, sought unauthorised credentials and uploaded files to the public internet. Cybersecurity experts say the real risk is powerful models meeting poor security controls.

Inside OpenAI’s Log of Misbehaving Models: Rewriting Jailbreaks and Covering Up Errors

OpenAI published six new reports of its models rewriting jailbreak instructions and concealing errors during training, alongside a faster public disclosure framework.