Guardrails
Screen every prompt before it reaches the agent, and audit what was blocked.
Where
Settings → Guardrails, and Settings → Guardrail Blocks for the audit trail.
What it does
Guardrails apply rules to every prompt on this instance, judged by a model kept deliberately separate from the one your employees run on: a jailbreak that fools the agent should not automatically fool the judge.

Two switches:
- Guardrail, screens every prompt before it reaches the agent.
- Also screen incoming messages, extends the check to email, Slack, Teams, Telegram, WhatsApp and webhooks.
A Raw tab shows the rules as text.
The default
Both switches are off. Nothing is screened on a fresh instance.
What breaks
A blocked incoming message is discarded silently, nobody is notified. Turning the second switch on with a judge model you do not trust means losing legitimate email without a trace. Turn it on knowingly.
Guardrail Blocks
The blocked prompts audit: what was blocked, when, by which rule, for which user, and the judge's reasoning.

The person who was blocked only ever sees a rule number, never the rule text, so this screen is the only place where the reason is legible.