Limits
Limits are applied per API key. Requests authenticated without a key fall back to a per-IP bucket (collapsed to a/64 for IPv6 clients).
Writes count against both buckets. A write rejected by the write bucket does
not consume global budget — one over-limit request produces exactly one
rejection.
These limits are generous by design and are not a billing control. Metered
usage — meeting minutes and LLM tokens — is governed by your plan, not by
these limits. Need a higher ceiling for a launch or a backfill? Contact us
before you need it, not during.
Headers
Present on every public API response, not just 429s.When you’re limited
A 429 returns the standard error envelope witherror.type of rate_limit_error:
Retry-After header over a fixed sleep — it reflects the
actual remaining window rather than a guess.
Staying under the limit
- Watch
RateLimit-Remainingand slow down before you hit zero, rather than sprinting into a 429 and backing off. - Prefer webhooks over polling. Nearly every polling loop we see is
waiting for a meeting to end or an analysis to land — both of which are
webhook events (
meeting.ended,session.analysis_complete). One subscription replaces thousands of requests. - Bucket by key, not by process. Limits are per API key, so ten workers sharing one key share one budget. Issue a key per workload if you want them isolated.

