Guardrails are prompt-level. They shape what the persona is instructed to
do; they are not a separate model that inspects and blocks each reply before it
is spoken. See What guardrails are not.
When to use a guardrail
- Keep the persona on-topic when the scenario prompt is broad
- Ban a subject outright — competitors, pricing, legal or medical advice
- Force a disclaimer, a required question, or a closing step
- Stop the persona promising anything you cannot deliver
- Apply one compliance rule across every scenario without editing each prompt
The Guardrail object
Single-object endpoints return the object directly. List endpoints wrap
their rows in
data, which is where cursor pagination
merges has_more and next_cursor:Attaching to a scenario
A guardrail does nothing until it is attached. Attach is idempotent, and re-attaching a disabled pair re-activates it.
Changes apply from the scenario’s next meeting. A call already in progress
keeps the prompt it started with.
What the persona actually receives
At meeting start we fetch the scenario’s active guardrails and append one block to the end of the system prompt — after the meeting script, after any tool or knowledge-base guidance, so nothing downstream can outrank it:- Precedence is explicit. A guardrail usually exists to narrow a scenario prompt that says something broader, so it has to win that conflict.
- The persona is told to stay quiet about the rules. An agent that announces “I’ve been told not to discuss competitors” reads as broken. It should simply change the subject.
guardrail_name never appears in the prompt — only the text.
Writing a good guardrail_prompt
Write a direct instruction to the persona, not a policy document.- Say what to do instead, not only what to avoid — a rule with no escape route makes the persona stall when the subject comes up.
- One rule per guardrail. Separate guardrails can be attached, ordered and disabled independently; a paragraph of six rules cannot.
- Keep it short. Guardrails ride on every turn of every meeting. Across all
attached guardrails we cap the block at 20,000 characters and drop the
overflow in
positionorder.
What guardrails are not
Guardrails steer the model through its prompt. They are not an independent enforcement layer, and a determined participant can still push a model off its instructions.- No separate model inspects each reply and blocks it before it is spoken
- Nothing terminates the call automatically when a rule is broken
- No per-violation event is recorded
Lifecycle
If the guardrail lookup fails at meeting start, the meeting proceeds without
guardrails rather than failing to connect. Attaching a guardrail can never stop
a call from starting.

