Breaking AI Defences: Best-of-N Jailbreaking and the Future of AI Security

Breaking AI Defences is one of those subjects that keeps surfacing in every security conversation. The concept of “jailbreaking” in AI refers to techniques that bypass safeguards in advanced AI models, enabling harmful outputs. A recent study introduced the Best-of-N (BoN) Jailbreaking method, an innovative approach to exploit these vulnerabilities across various input types, including text, vision, and audio.

BoN Jailbreaking works by repeatedly altering prompts with small changes, such as randomising text capitalisation or adjusting audio pitch, until a harmful response is elicited. This method has proven highly effective, achieving an attack success rate of up to 89% on state-of-the-art AI models, such as GPT-4 and Claude, when testing with 10,000 variations of a prompt.

Key findings:

  • Multi-modality effectiveness: BoN extends beyond text to bypass defences in vision and audio models, with similar success rates.
  • Scalability: The algorithm’s success improves predictably with more computational resources, following a power-law trend.
  • Adaptability: Combining BoN with other techniques like prefix optimisation significantly boosts its efficiency, reducing the number of attempts needed to bypass safeguards.

While this research highlights the ingenuity of attackers, it also underscores the pressing need for stronger, multi-modal defences in AI systems. Organisations relying on AI technologies must proactively evaluate and reinforce their security measures to stay ahead of these threats.

 

AI Tools
AI Tools

Read more from the original PDF here.

By understanding and addressing vulnerabilities like those exposed by BoN Jailbreaking, we can enhance the resilience and safety of AI technologies in an increasingly interconnected world.

If you have any related ideas, comments or views, feel free to share in the comments section below.

Related Reading

Subscribe

Related articles

NVIDIA, Microsoft, Meta, and 50+ Companies Tell Washington Not to Lock Down Open AI

More than 50 tech companies including NVIDIA, Microsoft, Meta, and Google published a joint letter urging Washington to protect open-weight AI models. The only notable holdout? Anthropic. Here is what the fight is actually about.

Europe Just Drew a Red Line on AI and Cybersecurity. Here Is What It Means

The EU Action Plan on Cybersecurity and AI sets nine actions across three pillars, with enforcement starting 2 August 2026. Fines climb to 3% of global turnover or 15 million euros for non-compliance. Here is what every security leader needs to know.

Australia Sets Rules for AI. The Hard Part Comes Next.

Australian writers, musicians and journalists will keep ownership of...

FLUX 3: How Black Forest Labs Is Bridging Video AI and Real-World Robots

Black Forest Labs' FLUX 3 is expanding from video and image generation into robot control for Audi factories, with an open-weight model planned for factory hardware.

The Open Source AI Revolution: When the World’s Biggest Models Became Free

Chinese open-source AI models led by Moonshot Kimi K3 have functionally closed the gap with proprietary systems from OpenAI and Anthropic, with profound consequences for geopolitics, global markets, and the future of autonomous AI agents.