Skip to main content
The Voice pill controls how an agent sounds. It is the last of the four pills at the top of the agent editor, and it configures text-to-speech (TTS) only - which languages the agent speaks is set in Language, and how it hears them in Transcriber. The pill shows the voice the agent is currently on. Open it and select Browse all voices for the full grid, with filters and previews.

Voices are filtered by the agent’s languages

The voice grid is narrowed by the languages you selected in the Language pill. A provider whose voices cannot cover them is greyed out, with the reason on hover.
Some providers cannot be handed some languages at all. Cartesia raises an error on Kannada, Odia, Assamese, and Urdu; Vocily AI’s catalogue is India-only. That is not a “sounds wrong” problem - the call fails to start, on every voice in that provider’s catalogue.If you change the primary language to one your current provider cannot speak, Vocily swaps the provider or the voice and says so in the pill. It never does this silently mid-call - see Language.

Choose a TTS provider

Five providers are available: Select a provider first, then a voice. Each provider currently offers a single model, so there is normally no model to choose - where a Model selector does appear, set it before picking a voice, because the voice list depends on it.
The voice catalogue changes often. The picker in your workspace is the live list - this page describes how choosing works, not which voices exist today.

Vocily AI

Vocily AI is Vocily’s own voice engine, first in the provider list and the one a new agent starts on. It is the lowest-cost provider in the catalog per minute, and it starts speaking in around 100ms. It speaks ten locales: Hindi, Indian English, Bengali, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu. For an agent whose languages fall outside that list - Odia, or any of the European or East Asian languages - a new agent starts on Sarvam AI or Cartesia instead, whichever covers the language.
Urdu and Assamese currently have no voice. They can be selected as agent languages and the speech models transcribe them, but no voice provider on the platform can speak them: Sarvam AI’s voices cover eleven languages and these are not among them, Cartesia refuses both, and Vocily AI’s catalogue is India-only without them.An agent with Urdu or Assamese as its primary language will not be able to speak. Use them only under Also speaks, where the agent can understand a caller who uses them and reply in a language it can voice.
US English and British English agents also start on Vocily AI, using its Indian English voice. Vocily AI has no American or British actor. If the accent matters for your callers, switch the provider to Cartesia, which offers English (US) and English (UK) voices.

The default voice

A new agent starts on a Vocily AI voice chosen for breadth of language coverage rather than for any one language - one certified for eight of the ten locales, so the agent’s voice stays valid if you change its language later. For the two locales that voice does not cover (Gujarati and Malayalam), a new agent starts on that locale’s own default voice instead. The pill always names the voice the agent is currently on.

Voices are certified per language

Unlike the other providers, a Vocily AI voice can only speak the languages it is certified for. Most cover one or two; a few cover many. Every voice card lists its own languages, so check the card before you choose.
A voice that cannot speak the agent’s language will save, but calls on it will not connect. Vocily never silently swaps a voice mid-call, so the mismatch is not repaired for you once the agent is live: it saves, and then every call fails to start. This is the one place Vocily AI differs from Cartesia, whose voices can read any supported language.Changing the primary language in the Language pill now repairs this for you at the moment you change it, and tells you it did. It is still worth a test call afterwards.
Place a test call after any language or voice change. A conversation that fails this way is recorded in Conversations with the status Failed.

Choose a voice

In Browse all voices, use the All, Female, or Male filters, the language filter, and the brand filter. Sarvam AI’s voices are marked multilingual, and on an agent with more than one language this is literal: the voice follows whichever language the agent replies in and changes mid-call by itself, with no setting to manage and no pause while it switches. Select Save Changes to keep the provider, model, and voice on the agent. For a phone agent, choose a voice that is clear at telephone quality, not only pleasant through a laptop speaker.

Speaking speed

Under the pill’s gear, Speaking speed runs from 0.5x to 1.5x. Start at 1.0x, then adjust after listening to a phone-quality preview.

Emotion

Cartesia is the only provider with an Emotion selector, shown in the voice grid. Alongside Default it offers Neutral, Curious, Excited, Enthusiastic, Happy, Content, Calm, Confident, Sad, Apologetic, and Frustrated. The control is hidden for every other provider.

Backup voice

Also under the gear: Backup voice, used for the rest of the call if the main voice provider stops responding. Select Choose a backup to open the same grid, so you can filter and preview the voice a caller would actually be handed - rather than picking one from a list nobody has ever listened to. Three things to know:
  • It must be a different brand. Failing over to the vendor that just stopped answering achieves nothing, so the primary provider is excluded from the backup grid.
  • It is a voice, not a brand. You cannot synthesize from a brand alone, so the backup is a specific voice.
  • Speed and emotion are not part of it. Those belong to the agent’s own voice.
Like the speech backup, a voice failover latches for the rest of the call and returns the conversation to the primary language, with the agent saying so. See Language.

Preview the voice

In Preview Voice, select Play to synthesize the current greeting. If the agent has a greeting, the preview uses it; otherwise Vocily uses a sample greeting. Choose a quality before listening:
  • Browser (HD) previews at browser-oriented quality.
  • Phone (16kHz) is a closer check of how the selected voice will translate to a call.
Previewing does not save the agent. Select Save Changes after you are happy with the provider, model, voice, speed, and emotion.

Voice quality checklist

Test the greeting and at least one longer response in both quality modes. Check that:
  • The first sentence is easy to understand without repeating it.
  • Names, numbers, email addresses, URLs, and IDs are pronounced acceptably.
  • The voice remains intelligible at the chosen speed.
  • The voice can speak every language in the agent’s set, not just the primary one.
  • The Transcriber turn-taking dials do not make the conversation feel rushed or difficult to interrupt.
The voice pipeline uses the selected speech model, LLM, and voice together. If the agent is not recognizing a language correctly, that is a Transcriber problem, not a voice one.