Skip to main content
Synthesize speech
Send text and a voice, get the audio back in the response body. Up to 1,000 characters a request. It is the same engine your agent speaks with, so it is the way to hear a voice before you put it on one — and the way to render a line your own product will play.
provider and voice_id both come from GET /v1/voices, and they travel together — a voice id means nothing without the provider it belongs to.

The audio

The response body is the audio — write it to a file or play it.
Nothing is stored. There is no id and no URL to fetch it again, so save the bytes if you need them.

Which providers take which controls

speed and emotion are not universal. A provider that does not honour one would drop it silently, so sending it is refused rather than ignored — you find out here, not by wondering why the audio sounds unchanged. emotion also takes one of a fixed set: neutral curious excited enthusiastic happy content calm confident sad apologetic frustrated. Anything else is refused. Send an unsupported control and you get a 400 naming the providers that do take it:
Each provider clamps speed to its own range, so the same 0.8 is not identical across two of them.

Language

language is the language the text will be spoken in, not the language the text is written in. Leave it out and the provider’s default applies. A voice must be certified for the language you ask for — languages on each row of GET /v1/voices is that list. A pair that is not certified is refused rather than substituted, so you hear the failure here instead of discovering it on a live call.

Authorizations

Authorization
string
header
required

Your API key as a Bearer token, e.g. Authorization: Bearer vk_….

Body

application/json

The body of POST /v1/speech.

text
string
required

What to say. Up to 1000 characters.

Required string length: 1 - 1000
provider
string
required

Voice provider, from GET /v1/voices.

voice_id
string
required

Which voice, from GET /v1/voices?provider=. Must belong to provider — a voice id means nothing without the engine it came from.

model
string | null

Read-only in practice. Each provider ships exactly one model and it is filled in from provider; GET /v1/tts-capabilities names which.

language
string | null

The language to speak text in, e.g. hi-IN. Defaults to the provider's own default. A voice not certified for it is refused rather than substituted — languages on each GET /v1/voices row is that list.

speed
number | null

Speaking rate. 1.0 is the voice's natural pace, and each provider clamps it to its own range. Not supported by elevenlabs — sent for that provider it is refused rather than ignored.

Required range: 0.25 <= x <= 4
emotion
string | null

Emotional colour: neutral curious excited enthusiastic happy content calm confident sad apologetic frustrated. Cartesia only — sent for any other provider it is refused, because nothing else would speak it.

Response

The spoken audio — the response body is the audio file itself.