MITRE ATLAS 2026.09 / technique reference
AML.T0054
LLM Jailbreak
MITRE source definition
Adversaries may induce a large language model (LLM) to ignore, circumvent, or override its safety/alignment behaviors and/or guardrails to elicit outputs the model is intended to withhold. Once jailbroken, the LLM may be used in unintended ways by the adversary. Jailbreaks may be achieved via adversarial prompting, or by modifying model weights or safety mechanisms.
Adversaries may attempt a jailbreak for Defense Evasion of the LLM's guidelines and guardrails itself to then reveal information (ex: [LLM Data Leakage](/techniques/AML.T0057), [Discover LLM System Information](/techniques/AML.T0069)) or generate harmful content (ex: [Generate Malicious Commands](/techniques/AML.T0102), [Spearphishing via Social Engineering LLM](/techniques/AML.T0052.000)). They may also jailbreak a model for Privilege Escalation to invoke tools or perform actions for their own purposes (ex: [AI Agent Tool Invocation](/techniques/AML.T0053)) or abuse the agent for a Command and Control channel (ex: [AI Agent](/techniques/AML.T0108)).
Adversaries use a variety of strategies to craft jailbreak prompts. Prompts may target specific models or model families and are iterated upon until successful. Model providers actively update their model guardrails to make them more resistant to jailbreak prompts as new prompts are developed. Common strategies [[jailbreak-guide]] include but are not limited to:
- Instruction override: Use phrasing that attempts to supersede prior constraints (e.g. "ignore previous instructions"). - Roleplay / persona switching: Instruct the LLM to adopt an identity or mode that allows unrestricted answers (e.g. "as a security researcher"). - Fictionalization and hypotheticals: Instruct the LLM to include disallowed content as part of a story, screenplay, or educational scenario. - Separate intent from content: request analysis, examples, templates, or edge cases, that implicitly contain disallowed content. - Multi-turn escalation / Crescendo: Utilize a sequence of prompts that start benign, establish trust, then gradually cross policy boundaries with incremental prompts. - Constrained output formats: Instruct the LLM to output to a strict schema or format (e.g. JSON, YAML, code, or tables). - Data structure injection: Use structured prompts (e.g. YAML, JSON, XML, etc.) to steer the LLM to produce structured outputs such as tool schemas or workflow fragments. This can be used by the adversary to call tools, pass dangerous inputs into legitimate tools, or hijack workflows.[[zenity-dsi]] - Obfuscation and transformation: Use encoding, transformations, translation, or euphemisms, (e.g., base64 encoding, "describe it in another language"). - Create a high priority objective: Frame compliance as necessary to fulfill the user's main task (e.g. "to complete the evaluation," "to follow the spec," "to follow safety guidelines"). - Affirmation: Appending affirmations such as "sure" to the end of prompts can help bypass refusals to generate malicious or otherwise undesired content.[[cybernews]]
Adversaries may also use algorithmic approaches to generating jailbreak prompts [[jailbreak-zoo]] [[jailbreak-survey]]. Algorithmic jailbreak generation allows for automated methods that discover jailbreaks at scale. Some approaches automate manual strategies [[autodan]] [[gptfuzzer]] [[crescendo]] [[echo-chamber]] while others optimize a string of tokens directly [[universal]] to produce nonsensical text. Both black-box (applicable to commercial models where the adversary has only query access to the model) and white-box (applicable in the open-source setting, where the adversary has full access to the model weights) optimization approaches are viable.
Adversaries may also directly manipulate a model's weights, or modify or remove parts of a model to create a jailbroken or "uncensored" variant of the target model. This is applicable to open-source models, or cases where the adversary gains full access to the target model. Approaches include fine-tuning to reduce refusals [[single-direction]], targeted model editing [[rome]], addition of adapters [[lora]], and removing safety mechanisms such as guardrails.
Jailbreak prompts that are known to work on various classes of LLMs are often published in the open-source community [[dan]]. Jailbroken or uncensored LLMs that have been trained or fine-tuned to be jailbroken are shared in public model registries such as huggingface [[abliteration]].
Source modified 2026-09-15. Reproduced from the pinned ATLAS release; inline technique links resolve to local reference pages.
Parent, sub-techniques and ATT&CK references
No explicit relationship in this pinned source.
Source-backed defensive context
MITRE mitigations
- AML.M0020 Generative AI Guardrails
- AML.M0021 Generative AI Guidelines
- AML.M0022 Generative AI Model Alignment
- AML.M0035 AI Red Team
MITRE case studies
- AML.CS0041 Rules File Backdoor: Supply Chain Attack on AI Coding Assistants · Exercise
- AML.CS0046 Data Destruction via Indirect Prompt Injection Targeting Claude Computer-Use · Exercise
- AML.CS0051 OpenClaw Command & Control via Prompt Injection · Exercise
- AML.CS0052 LLMSmith: RCE Vulnerabilities in LLM-Integrated Applications · Exercise
- AML.CS0057 Storm-2139 Azure OpenAI Guardrail Bypass · Incident
- AML.CS0063 Prompt-Based Attacks Against Gemini via Calendar Invitations · Exercise
- AML.CS0066 ZombieAgent: Data Exfiltration Attack on ChatGPT · Exercise
- AML.CS0067 Claude Code GitHub Action Secret Exposure · Exercise
- AML.CS0069 GTG-1002 Claude Code Espionage Campaign · Incident
These are explicit source relationships, not independently reproduced incidents or validated detection coverage.
Simulation and telemetry boundary
This is a technique reference page, not a runnable simulation. No ATLAS-specific telemetry mapping, local attack execution or detector validation is asserted. MITRE maturity describes its source evidence, not a 1200km lab result.
For broader context—not technique-specific control mappings—see AI Security, AI Security Course, and detection-validation methodology.
Provenance and attribution
Immutable MITRE ATLAS source · Import provenance · Attribution and transformation notice · Apache License 2.0
Copyright 2021-2026 MITRE. Source text and explicit relationships are retained; navigation, formatting and local links are provided by 1200km.
- Uncensor any LLM with abliteration
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
- ChatGPT DAN
- The Echo Chamber Multi-Turn LLM Jailbreak
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Jailbreaking LLMs: A Comprehensive Guide (With Examples)
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
- JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- Locating and Editing Factual Associations in GPT
- Refusal in Language Models Is Mediated by a Single Direction
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- GitHub Copilot Jailbreak Vulnerability Let Attackers Train Malicious Models
- Data-Structure Injection (DSI) in AI Agents