GPT · Prompt injection · Decrypt
In the paper “ Prompt Injection as Role Confusion,” presented at the International Conference on Machine Learning in June
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
“For an LLM, everything arrives through the same channel as one long token soup,” the team wrote.
Key facts
- The researchers said the technique increased jailbreak success rates from near zero to about 60% across the models they tested, including OpenAI's GPT-5 nano, mini, and full, o4-mini, and gpt-oss-20b
- In June, Microsoft disclosed a prompt injection vulnerability in Anthropic's Claude Code GitHub Action that could have exposed credentials stored in software development pipelines
- Forget clever prompts: AI researchers say they tricked leading AI models into generating cocaine synthesis instructions by convincing them the dangerous ideas were their own, while also manipulating
- Called Chain-of-Thought (CoT) Forgery, the attack inserts fake reasoning that mimics a model's internal thought process
Summary
Researchers got frontier AI models to generate cocaine synthesis instructions using a new prompt injection attack. The same technique manipulated an AI coding agent into uploading sensitive credentials. The study argues prompt injection stems from "role confusion," not simply models failing to recognize malicious prompts. Forget clever prompts: AI researchers say they tricked leading AI models into generating cocaine synthesis instructions by convincing them the dangerous ideas were their own, while also manipulating an AI coding agent into leaking sensitive credentials. In the paper “ Prompt Injection as Role Confusion,” presented at the International Conference on Machine Learning in June, researchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell argue that both prompt injection attack demonstrations stem from a structural flaw in how large language models (LLMs) distinguish trusted instructions from untrusted text.