Skip to main content
Every response from the public API carries rate-limit headers, so you never have to discover a limit by hitting it.

Limits

Limits are applied per API key. Requests authenticated without a key fall back to a per-IP bucket (collapsed to a /64 for IPv6 clients). Writes count against both buckets. A write rejected by the write bucket does not consume global budget — one over-limit request produces exactly one rejection.
These limits are generous by design and are not a billing control. Metered usage — meeting minutes and LLM tokens — is governed by your plan, not by these limits. Need a higher ceiling for a launch or a backfill? Contact us before you need it, not during.

Headers

Present on every public API response, not just 429s.

When you’re limited

A 429 returns the standard error envelope with error.type of rate_limit_error:
Always prefer the Retry-After header over a fixed sleep — it reflects the actual remaining window rather than a guess.

Staying under the limit

  • Watch RateLimit-Remaining and slow down before you hit zero, rather than sprinting into a 429 and backing off.
  • Prefer webhooks over polling. Nearly every polling loop we see is waiting for a meeting to end or an analysis to land — both of which are webhook events (meeting.ended, session.analysis_complete). One subscription replaces thousands of requests.
  • Bucket by key, not by process. Limits are per API key, so ten workers sharing one key share one budget. Issue a key per workload if you want them isolated.