Aguardic logoAguardic

Prompt Injection & Jailbreak Prevention

by AguardicOfficial·v1.0.0

Detect and block prompt injection attacks and jailbreak attempts targeting AI models.

About This Policy Template

Critical AI safety policy for any organization exposing AI models to user input. Detects prompt injection patterns including system prompt extraction attempts, role-play jailbreaks, instruction override attacks, encoding-based bypasses, and common jailbreak templates. Essential for customer-facing AI chatbots, copilots, and agent systems.

Policy Rules(3)

High Severity

(3)

Instruction Override Attempt

Detect attempts to override safety instructions via injected context

Rule

Role-Play Jailbreak Detection

Detect jailbreak attempts using role-play or persona switching

Rule

System Prompt Extraction Attempt

Detect attempts to extract or reveal the system prompt

Rule

Enforcement by Integration

What happens when a violation is detected, based on the enforcement mode and integration type.

IntegrationBlockApprovalWarnMonitor
Version Control
GitHub, GitLab, Bitbucket
Fail check run / merge request statusPending check run, held for reviewNeutral check run / comment on PRPass check run (silent)
Email · Gmail
Gmail
Quarantine label; + violation label (outbound)Quarantine label, held for reviewAdd warning labelLog only
Email · Outlook
Outlook
Move to quarantine folder; + flag (outbound)Move to quarantine, held for reviewFlag + categorizeLog only
Messaging
Slack, Teams
Post violation warning in channelPost 'held for review' warningPost warning in channelLog only
Storage
Google Drive, Dropbox, OneDrive
Move file to quarantine folderQuarantine file, held for reviewLog onlyLog only
AI Proxy
OpenAI, Anthropic, Gemini, MCP, Agent
Block request (return 403)Hold request, return review IDAllow request + audit trailLog only
API
REST API
Return BLOCK outcome (client decides)Return APPROVAL_REQUIRED + poll URLReturn WARN outcomeLog only

Version History

1 version published

v1.0.0Active2/23/2026

Initial release

Got a vendor security questionnaire?

Answer the AI questions with controls Aguardic enforces

Try the tool

Ready to Install Prompt Injection & Jailbreak Prevention?

Get started with pre-built governance policies in minutes.