Prompt Injection & Jailbreak Prevention
Detect and block prompt injection attacks and jailbreak attempts targeting AI models.
About This Policy Template
Critical AI safety policy for any organization exposing AI models to user input. Detects prompt injection patterns including system prompt extraction attempts, role-play jailbreaks, instruction override attacks, encoding-based bypasses, and common jailbreak templates. Essential for customer-facing AI chatbots, copilots, and agent systems.
Policy Rules(3)
High Severity
(3)Instruction Override Attempt
Detect attempts to override safety instructions via injected context
Role-Play Jailbreak Detection
Detect jailbreak attempts using role-play or persona switching
System Prompt Extraction Attempt
Detect attempts to extract or reveal the system prompt
Enforcement by Integration
What happens when a violation is detected, based on the enforcement mode and integration type.
| Integration | Block | Approval | Warn | Monitor |
|---|---|---|---|---|
Version Control GitHub, GitLab, Bitbucket | Fail check run / merge request status | Pending check run, held for review | Neutral check run / comment on PR | Pass check run (silent) |
Email · Gmail Gmail | Quarantine label; + violation label (outbound) | Quarantine label, held for review | Add warning label | Log only |
Email · Outlook Outlook | Move to quarantine folder; + flag (outbound) | Move to quarantine, held for review | Flag + categorize | Log only |
Messaging Slack, Teams | Post violation warning in channel | Post 'held for review' warning | Post warning in channel | Log only |
Storage Google Drive, Dropbox, OneDrive | Move file to quarantine folder | Quarantine file, held for review | Log only | Log only |
AI Proxy OpenAI, Anthropic, Gemini, MCP, Agent | Block request (return 403) | Hold request, return review ID | Allow request + audit trail | Log only |
API REST API | Return BLOCK outcome (client decides) | Return APPROVAL_REQUIRED + poll URL | Return WARN outcome | Log only |
Version History
1 version published
Initial release
Got a vendor security questionnaire?
Answer the AI questions with controls Aguardic enforces
Ready to Install Prompt Injection & Jailbreak Prevention?
Get started with pre-built governance policies in minutes.