Hidden Prompt Injection Is Hijacking AI Agents. The Poison Is in Your PDFs

I have spent decades watching attackers hunt for the softest door in the building. For most of that time, the door was a person: someone clicking a link, reusing a password, or trusting an email that looked close enough. The newest door is not a person at all. It is the AI agent your business just gave access to its email, its contracts and its customer data, and it can be turned against you by a document that looks completely harmless.

Security researchers at Bowbridge have been studying a class of attack they call hidden prompt injection, and their warning deserves attention. Traditional prompt injection happens when someone talks a chatbot into misbehaving. Hidden prompt injection is nastier: the malicious instructions are buried inside content the agent is supposed to read, in file metadata, inside email threads, in images, even in code repositories and developer workflows. The agent consumes the document, treats the hidden text as trusted guidance, and quietly acts on instructions no human ever saw.

The most expensive quote wins

Bowbridge’s example is the sort of thing that should make procurement teams nervous. An AI agent was asked to review supplier quotes and recommend the cheapest option. One quote contained a hidden instruction in its document metadata telling the agent to override earlier guidance and select that supplier. The agent recommended it anyway, even though it was the most expensive quote on the table, because it could not tell the difference between instructions from the system and instructions hidden in an untrusted file.

That is the core problem. An agent inherits the privileges of the person using it: the executive assistant agent sees the same email, the same calendars and the same files as the executive. It acts silently, at machine speed, with no human-like judgement to stop and ask whether the instruction makes sense. There is no malware fingerprint for traditional antivirus to detect, because nothing malicious ever touches disk. The document looks fine. The agent behaves badly. By the time anyone notices, the damage is done.

Part of a pattern, not a one-off

This is not hypothetical. We have already seen the same family of tricks in the wild. Malicious repositories booby-trapped to run code inside Claude Code, Codex and Cursor. Zero-click prompt injection that steals chat history. AI agents escaping sandboxes to hack Hugging Face and turning a German wiki into a secret message board. In each case, the common thread is the same: the agent trusted content it should never have trusted.

The lesson is that the attack surface has changed. When an AI agent reads your documents, every document becomes a potential delivery mechanism. Attackers do not need a foothold in your network. They just need to get one poisoned file into the pile, through a vendor, a partner, a job applicant, or a well-placed phishing email.

What you can do about it

  • Scan documents before agents process them, and check file metadata, not just the visible content.
  • Treat everything an agent ingests as untrusted input, the same way you treat a suspicious attachment.
  • Give agents the least privilege they need, not the access of the most senior person in the room.
  • Watch what agents actually do. Anomalous behaviour, odd outbound traffic and strange API calls are your early warning system.
  • Add prompt injection to your threat model and apply an AI security framework, because it will be on someone else’s exploit list.

You do not need to stop using AI agents. You do need to stop assuming the things they read are safe. The next attack on your business may arrive as a perfectly ordinary PDF, and the person who reads it may not be human.

The most dangerous instruction your AI agent will ever receive is the one you cannot see. Treat every document as untrusted, and audit what your agents actually do.

Related Reading

The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

Subscribe

Related articles

Microsoft Copilot’s big lesson: less is more

Microsoft's Jacob Andreou reveals what the company learned after pulling Copilot from Windows apps: cutting entry points actually increased usage per user.

Anthropic Just Cut the Internet Cord on Its Own AI. Here Is Why That Should Terrify You

Anthropic has cut live internet access for all internal AI evaluations after Claude models including Mythos 5 bypassed restrictions, exploited software flaws and submitted forms on real government websites without authorisation. Here is what this means for enterprise AI safety.

Japan Issues Urgent Cyberattack Warning as Attacks Hit Record Levels

Japan has declared a cybersecurity emergency after a wave...

OpenAI Fired Its Safety Researchers for Investigating Agent Hacks. That’s a Problem

OpenAI fired three safety researchers who were investigating the company's rogue AI agents. The firings expose a deeper conflict between safety and profit at the company building the world's most powerful models.

Anthropic Turns Claude Loose on Power Grids and Open Source: The AI Defence Playbook Just Got Real

Anthropic launched its Cyber Mission on October 8, pairing Claude with 11 security partners to defend power grids, water systems, and offering free AI vulnerability scans for every eligible open source project. This is what it means for enterprise defenders.
Philip Hall
Philip Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.