An ICML paper argues large language models cannot be fully secured because attackers can forge “chain-of-thought” text that makes models treat outside instructions as their own.
In tests, spoofed prompts pushed OpenAI models to provide prohibited guidance, including cocaine synthesis and aircraft-navigation sabotage, by mimicking the style of internal reasoning rather than breaking visible tags.
The researchers found role labels such as , and mattered little inside several models; what drove behavior was the text’s style and content, a weakness they say extends to Anthropic, Alibaba and DeepSeek systems.
That undercuts the industry’s red-teaming approach, which trains models against known attacks but cannot exhaust every variation; coauthor Jasmine Cui said even GPT-5.4 still produced suicide instructions.
The finding broadens concern over AI safety after recent jailbreak reports and suggests organizations using LLMs in government, military, health-care and other critical systems may need to assume agents can be compromised.