> ## Documentation Index
> Fetch the complete documentation index at: https://docs.waterr.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Bring Your Own Agent

> Point your OpenAI-compatible agent at Waterr and it joins a real-time video meeting — no SDK, no rewrite, your infra.

If you already have an agent — a loop over chat completions with its own tools,
state, and guardrails — you can make it the **brain of a live Waterr meeting**.
Waterr provides the video room, transport, speech-to-text, text-to-speech,
turn-taking, transcripts, and post-meeting analysis. Your endpoint provides the
words.

Your agent loop, tools, and state stay entirely on your infrastructure. From
Waterr's side, your agent is just a streaming model.

## How it works

1. You expose an **OpenAI-compatible streaming chat-completions endpoint**
   (`POST {base_url}/chat/completions` with `stream: true`).
2. You attach it to a scenario (`custom_agent_config`) or pass it per-meeting.
3. When a meeting starts, Waterr's pipeline transcribes the participant,
   calls your endpoint with the running conversation, and speaks your streamed
   reply into the room.

<Note>
  Bring Your Own Agent runs on the cascaded (STT→LLM→TTS) pipeline. If the
  scenario also has a speech-native mode enabled (live vision / realtime
  backends), the custom agent wins and the meeting uses the cascaded path.
</Note>

## Configuration

Attach the connector to a scenario via the standard scenario API:

```bash theme={null}
curl -X PUT https://api.waterr.ai/v1/scenarios/{scenario_id} \
  -H "Authorization: Bearer wai_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "custom_agent_enabled": true,
    "custom_agent_config": {
      "base_url": "https://agent.yourcompany.com/v1",
      "api_key": "sk-your-endpoint-key",
      "model": "support-bot-v3",
      "timeout_s": 10,
      "extra_headers": {"X-Tenant": "acme"}
    }
  }'
```

| Field | Required | Description |
| - | - | - |
| `base_url` | ✓ | HTTPS URL of your OpenAI-compatible API root (the part before `/chat/completions`). |
| `api_key` | – | Sent to your endpoint as `Authorization: Bearer …`. Stored encrypted; masked (`sk-…abc4`) on every read. |
| `model` | – | Opaque string passed through in the request body — route on it server-side. Default `waterr-agent`. |
| `timeout_s` | – | First-response budget per completion, capped at 30. Default 10. |
| `extra_headers` | – | Additional headers. `Host`, `Authorization`, and `X-Internal-*` are rejected. |

Every meeting created on that scenario — API, Studio, or embed — then uses
your agent.

## Endpoint contract

Your endpoint must:

* **Accept** `POST {base_url}/chat/completions` with an OpenAI-shaped body:
  `model`, `messages` (the running meeting conversation, system prompt
  included), optional `tools`.
* **Stream** the reply as SSE `chat.completion.chunk` deltas
  (`choices[0].delta.content`), ending with `data: [DONE]`.
* **Answer fast.** This is a live conversation: target **first token ≤ 1.5s**.
  Waterr masks slow turns with filler audio and retries once after
  `timeout_s`, but a sluggish endpoint ruins the conversational feel — that
  part of the latency budget is yours.

Any framework that can serve an OpenAI-compatible endpoint works: vLLM,
LiteLLM, a LangServe route, or a \~40-line wrapper around your existing agent:

```python theme={null}
# pip install fastapi uvicorn
import json, time
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse

app = FastAPI()

@app.post("/v1/chat/completions")
async def chat(request: Request):
    body = await request.json()
    reply_stream = my_agent.run(body["messages"])   # your existing agent loop

    def sse():
        for token in reply_stream:
            chunk = {
                "id": "chatcmpl-waterr", "object": "chat.completion.chunk",
                "created": int(time.time()), "model": body.get("model"),
                "choices": [{"index": 0, "delta": {"content": token},
                             "finish_reason": None}],
            }
            yield f"data: {json.dumps(chunk)}\n\n"
        yield 'data: {"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}\n\n'
        yield "data: [DONE]\n\n"

    return StreamingResponse(sse(), media_type="text/event-stream")
```

## Platform tools

Waterr sends its built-in function schemas (`end_call`, web search, your
scenario's [custom functions](/api-reference/custom-functions)) in the
standard `tools` field. Your endpoint may honor them by emitting standard
`tool_calls` deltas — or ignore them entirely. Your **own** tools run inside
your agent loop, behind your endpoint, invisible to Waterr. That's the point.

## Security

* Waterr never sends any internal credential to your endpoint — only the
  `api_key` you configured.
* `base_url` must be public HTTPS. URLs resolving to private, loopback, or
  link-local address space are rejected at meeting creation.
* Your `api_key` is encrypted at rest (AES-256-GCM) and masked in every API
  response.
* Embed meetings only ever use the scenario's stored config — an embedded
  page can never inject its own agent endpoint.

## Billing

Meetings consume session minutes as usual (Waterr still runs transport, STT,
TTS, and analysis). LLM inference happens on your infrastructure, so **no LLM
token usage is metered** for custom-agent meetings.

## Close the loop

Pair the connector with [webhooks](/api-reference/webhooks): your agent takes
the meeting, then `transcript.ready` / `session.analysis_complete` deliver the
transcript and goal-scored analysis back to your system — a full loop where
Waterr is the face and voice, and your stack is the brain and memory.

Building with LangChain? The [LangChain integration](/api-reference/langchain)
covers both directions — agents that *manage* meetings and agents that *take*
them.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.