How it works
- You expose an OpenAI-compatible streaming chat-completions endpoint
(
POST {base_url}/chat/completionswithstream: true). - You attach it to a scenario (
custom_agent_config) or pass it per-meeting. - When a meeting starts, Waterr’s pipeline transcribes the participant, calls your endpoint with the running conversation, and speaks your streamed reply into the room.
Bring Your Own Agent runs on the cascaded (STT→LLM→TTS) pipeline. If the
scenario also has a speech-native mode enabled (live vision / realtime
backends), the custom agent wins and the meeting uses the cascaded path.
Configuration
Attach the connector to a scenario via the standard scenario API:
Every meeting created on that scenario — API, Studio, or embed — then uses
your agent.
Endpoint contract
Your endpoint must:- Accept
POST {base_url}/chat/completionswith an OpenAI-shaped body:model,messages(the running meeting conversation, system prompt included), optionaltools. - Stream the reply as SSE
chat.completion.chunkdeltas (choices[0].delta.content), ending withdata: [DONE]. - Answer fast. This is a live conversation: target first token ≤ 1.5s.
Waterr masks slow turns with filler audio and retries once after
timeout_s, but a sluggish endpoint ruins the conversational feel — that part of the latency budget is yours.
Platform tools
Waterr sends its built-in function schemas (end_call, web search, your
scenario’s custom functions) in the
standard tools field. Your endpoint may honor them by emitting standard
tool_calls deltas — or ignore them entirely. Your own tools run inside
your agent loop, behind your endpoint, invisible to Waterr. That’s the point.
Security
- Waterr never sends any internal credential to your endpoint — only the
api_keyyou configured. base_urlmust be public HTTPS. URLs resolving to private, loopback, or link-local address space are rejected at meeting creation.- Your
api_keyis encrypted at rest (AES-256-GCM) and masked in every API response. - Embed meetings only ever use the scenario’s stored config — an embedded page can never inject its own agent endpoint.
Billing
Meetings consume session minutes as usual (Waterr still runs transport, STT, TTS, and analysis). LLM inference happens on your infrastructure, so no LLM token usage is metered for custom-agent meetings.Close the loop
Pair the connector with webhooks: your agent takes the meeting, thentranscript.ready / session.analysis_complete deliver the
transcript and goal-scored analysis back to your system — a full loop where
Waterr is the face and voice, and your stack is the brain and memory.
Building with LangChain? The LangChain integration
covers both directions — agents that manage meetings and agents that take
them.
