Speech foundations
ORI-Realtime 1.5 builds on open speech models released by Kyutai:- Kyutai streaming TTS (kyutai.org/tts) — text-to-speech model weights, CC BY 4.0.
- Kyutai streaming STT — speech-to-text model weights, CC BY 4.0.
- Unmute serving components — open-source code released by Kyutai under permissive licenses (MIT/Apache-2.0).
Reference voices
Waterr’s voice presets use reference audio from the following datasets. Theid column matches the voice_id returned by GET /voices.
CSTR VCTK Corpus
The voices below are derived from the CSTR VCTK Corpus (Centre for Speech Technology Research, University of Edinburgh), released under the CC BY 4.0 license.
Citation: Yamagishi, Junichi; Veaux, Christophe; MacDonald, Kirsten. (2019). CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92). University of Edinburgh. The Centre for Speech Technology Research (CSTR). https://doi.org/10.7488/ds/2645
License summary
If you build a downstream product on top of Waterr’s voice API, these attributions must be surfaced to your users in a comparable location (e.g. your own docs or an About page). CC BY 4.0 permits commercial use with attribution; it does not permit removing the attribution requirement.

