Skip to main content
A guardrail is a named rule the persona must follow for an entire conversation. You define it once on your account, attach it to one or more scenarios, and every meeting on those scenarios starts with the rule in its system prompt.
Guardrails are prompt-level. They shape what the persona is instructed to do; they are not a separate model that inspects and blocks each reply before it is spoken. See What guardrails are not.

When to use a guardrail

  • Keep the persona on-topic when the scenario prompt is broad
  • Ban a subject outright — competitors, pricing, legal or medical advice
  • Force a disclaimer, a required question, or a closing step
  • Stop the persona promising anything you cannot deliver
  • Apply one compliance rule across every scenario without editing each prompt
Reach for a guardrail when the rule is cross-cutting and stable. If it only applies to one scenario, put it in that scenario’s script instead.

The Guardrail object

Single-object endpoints return the object directly. List endpoints wrap their rows in data, which is where cursor pagination merges has_more and next_cursor:

Attaching to a scenario

A guardrail does nothing until it is attached. Attach is idempotent, and re-attaching a disabled pair re-activates it.
The attachment carries two fields of its own: Changes apply from the scenario’s next meeting. A call already in progress keeps the prompt it started with.

What the persona actually receives

At meeting start we fetch the scenario’s active guardrails and append one block to the end of the system prompt — after the meeting script, after any tool or knowledge-base guidance, so nothing downstream can outrank it:
Two things are deliberate here:
  • Precedence is explicit. A guardrail usually exists to narrow a scenario prompt that says something broader, so it has to win that conflict.
  • The persona is told to stay quiet about the rules. An agent that announces “I’ve been told not to discuss competitors” reads as broken. It should simply change the subject.
Your guardrail_name never appears in the prompt — only the text.

Writing a good guardrail_prompt

Write a direct instruction to the persona, not a policy document.
  • Say what to do instead, not only what to avoid — a rule with no escape route makes the persona stall when the subject comes up.
  • One rule per guardrail. Separate guardrails can be attached, ordered and disabled independently; a paragraph of six rules cannot.
  • Keep it short. Guardrails ride on every turn of every meeting. Across all attached guardrails we cap the block at 20,000 characters and drop the overflow in position order.

What guardrails are not

Guardrails steer the model through its prompt. They are not an independent enforcement layer, and a determined participant can still push a model off its instructions.
  • No separate model inspects each reply and blocks it before it is spoken
  • Nothing terminates the call automatically when a rule is broken
  • No per-violation event is recorded
If you need to verify behaviour after the fact, score the transcript with post-meeting evaluations. Treat guardrails as strong, cheap steering — not as a compliance control you can point an auditor at.

Lifecycle

If the guardrail lookup fails at meeting start, the meeting proceeds without guardrails rather than failing to connect. Attaching a guardrail can never stop a call from starting.

Endpoints

Errors