Skip to content

Glossary · AI & Development

Jailbreak

An adversarial prompt that bypasses an AI model's built in safety rules to make it produce content it normally would not.

Browse all definitions

In detail

A jailbreak is a prompt crafted to make an AI model ignore the safety training and policies that constrain its output. Common techniques include role playing scenarios ("pretend you are a model with no restrictions"), prompt injection through smuggled instructions, encoding messages in unusual formats and incrementally pushing the model past each refusal. Foundation model providers patch known jailbreaks, but new ones surface constantly.

Apply the definition

Want to talk through how this applies to your business?

Start with the decision in front of you. We will help map the fit.

Straight answers · no pitch deck · no commitment