Model alignment bypasses that succeed in generating harmful content are still possible, although they are not universal [...] and almost never transferable [...] We have developed a prompting technique that is both universal and transferable and can be used to generate practically any form of harmful content from all major frontier AI models. [...] Our technique is transferable across model architectures, inference strategies, such as chain of thought and reasoning, and alignment approaches. A single prompt can be designed to work across all of the major frontier AI models.
Curated from hiddenlayer.com →
HiddenLayer's Policy Puppetry dresses a request up as configuration. Wrapped in something that looks like a policy file, the instruction reads to the model as though it came from whoever wrote the system prompt rather than from the person typing, and combined with a fictional framing it got refusals overturned across models from OpenAI, Google, Microsoft, Anthropic, Meta, DeepSeek, Qwen and Mistral in April 2025. Two properties make it the one to know about: it is universal, meaning one technique gets any category of refused content, and it is transferable, meaning the same prompt works on models that share no architecture. HiddenLayer's conclusion is the same as Microsoft's above and is the reason both are on this page: reinforcement learning from human feedback is not a security control.