Synthesize speech
Turn text into audio with any voice from the catalogue.
provider and voice_id both come from GET /v1/voices, and they
travel together — a voice id means nothing without the provider it belongs to.
The audio
The response body is the audio — write it to a file or play it.Which providers take which controls
speed and emotion are not universal. A provider that does not honour one would drop it
silently, so sending it is refused rather than ignored — you find out here, not by wondering
why the audio sounds unchanged.
emotion also takes one of a fixed set: neutral curious excited enthusiastic happy
content calm confident sad apologetic frustrated. Anything else is refused.
Send an unsupported control and you get a 400 naming the providers that do take it:
speed to its own range, so the same 0.8 is not identical across two of
them.
Language
language is the language the text will be spoken in, not the language the text is written in.
Leave it out and the provider’s default applies.
A voice must be certified for the language you ask for — languages on each row of
GET /v1/voices is that list. A pair that is not certified is
refused rather than substituted, so you hear the failure here instead of discovering it on a live
call.Authorizations
Your API key as a Bearer token, e.g. Authorization: Bearer vk_….
Body
The body of POST /v1/speech.
What to say. Up to 1000 characters.
1 - 1000Voice provider, from GET /v1/voices.
Which voice, from GET /v1/voices?provider=. Must belong to provider — a voice id means nothing without the engine it came from.
Read-only in practice. Each provider ships exactly one model and it is filled in from provider; GET /v1/tts-capabilities names which.
The language to speak text in, e.g. hi-IN. Defaults to the provider's own default. A voice not certified for it is refused rather than substituted — languages on each GET /v1/voices row is that list.
Speaking rate. 1.0 is the voice's natural pace, and each provider clamps it to its own range. Not supported by elevenlabs — sent for that provider it is refused rather than ignored.
0.25 <= x <= 4Emotional colour: neutral curious excited enthusiastic happy content calm confident sad apologetic frustrated. Cartesia only — sent for any other provider it is refused, because nothing else would speak it.
Response
The spoken audio — the response body is the audio file itself.