Jailbreak AI
Noun · AI & Machine Learning
Definitions
An attempt to bypass an AI system's intended safeguards, restrictions, or instruction boundaries through adversarial prompting or workflow manipulation. Jailbreak attempts are a major concern in public-facing AI systems because they can expose unsafe or disallowed behavior.
In plain English: A way of trying to get an AI system to ignore its safety rules.
Example: "The red team found a jailbreak AI pattern that tricked the assistant into treating hostile instructions as fictional role-play."