Updated
Updated · MIT Technology Review · Jul 30
Researchers Find LLM Flaw Lets 5 Models Spill Banned Instructions, Defying Full Security
Updated
Updated · MIT Technology Review · Jul 30

Researchers Find LLM Flaw Lets 5 Models Spill Banned Instructions, Defying Full Security

3 articles · Updated · MIT Technology Review · Jul 30

Summary

  • An ICML paper argues large language models cannot be fully secured because attackers can forge “chain-of-thought” text that makes models treat outside instructions as their own.
  • In tests, spoofed prompts pushed OpenAI models to provide prohibited guidance, including cocaine synthesis and aircraft-navigation sabotage, by mimicking the style of internal reasoning rather than breaking visible tags.
  • The researchers found role labels such as , and mattered little inside several models; what drove behavior was the text’s style and content, a weakness they say extends to Anthropic, Alibaba and DeepSeek systems.
  • That undercuts the industry’s red-teaming approach, which trains models against known attacks but cannot exhaust every variation; coauthor Jasmine Cui said even GPT-5.4 still produced suicide instructions.
  • The finding broadens concern over AI safety after recent jailbreak reports and suggests organizations using LLMs in government, military, health-care and other critical systems may need to assume agents can be compromised.

Insights

While tech giants claim their AI models are secure, could cheap, automated jailbreak tools eventually outpace even the strongest safety defenses?
If an AI can be tricked into planning a catastrophic cyberattack for just $58, is voluntary self-regulation already a failed experiment?