> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vocily.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Synthesize speech

> Turn text into audio with any voice from the catalogue.

Send text and a voice, get the audio back in the response body. Up to 1,000 characters a request.

It is the same engine your agent speaks with, so it is the way to hear a voice before you put it on
one — and the way to render a line your own product will play.

```http theme={"dark"}
POST /v1/speech
Content-Type: application/json

{
  "text": "Hi, this is Riya calling from Acme about your appointment tomorrow.",
  "provider": "cartesia",
  "voice_id": "f6141af3-5f94-418c-80ed-a45d450e7e2e",
  "language": "en"
}
```

`provider` and `voice_id` both come from [`GET /v1/voices`](/developers/catalogues/voices), and they
travel together — a voice id means nothing without the provider it belongs to.

## The audio

The response body **is** the audio — write it to a file or play it.

```bash theme={"dark"}
curl -X POST https://api.vocily.ai/v1/speech \
  -H "Authorization: Bearer vk_..." \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello there.","provider":"cartesia","voice_id":"f6141af3-5f94-418c-80ed-a45d450e7e2e"}' \
  --output hello.mp3
```

Nothing is stored. There is no id and no URL to fetch it again, so save the bytes if you need them.

## Which providers take which controls

`speed` and `emotion` are not universal. A provider that does not honour one would drop it
silently, so sending it is **refused** rather than ignored — you find out here, not by wondering
why the audio sounds unchanged.

| Provider     | `speed` | `emotion` |
| ------------ | :-----: | :-------: |
| `cartesia`   |   yes   |    yes    |
| `vocily`     |   yes   |     —     |
| `sarvam`     |   yes   |     —     |
| `smallest`   |   yes   |     —     |
| `elevenlabs` |    —    |     —     |

`emotion` also takes one of a fixed set: `neutral` `curious` `excited` `enthusiastic` `happy`
`content` `calm` `confident` `sad` `apologetic` `frustrated`. Anything else is refused.

Send an unsupported control and you get a `400` naming the providers that do take it:

```json theme={"dark"}
{
  "detail": "'speed' is not supported by 'elevenlabs' and would be ignored. It is honoured by: cartesia, sarvam, smallest, vocily. Remove it, or use one of those providers.",
  "code": "BAD_REQUEST"
}
```

Each provider clamps `speed` to its own range, so the same `0.8` is not identical across two of
them.

## Language

`language` is the language the text will be spoken in, not the language the text is written in.
Leave it out and the provider's default applies.

A voice must be certified for the language you ask for — `languages` on each row of
[`GET /v1/voices`](/developers/catalogues/voices) is that list. A pair that is not certified is
refused rather than substituted, so you hear the failure here instead of discovering it on a live
call.


## OpenAPI

````yaml developers/openapi.json POST /v1/speech
openapi: 3.1.0
info:
  title: Vocily API
  description: >-
    Public REST API for Vocily. Build and configure an agent, publish a version
    and put it live, place outbound calls, and read back calls, chats and what
    the agent remembered. Authenticate with a workspace API key as a Bearer
    token.


    Some things stay in the dashboard, by design: creating an API key, buying or
    connecting a phone number, setting an agent's webhook URL, connecting
    WhatsApp and its templates, building HTTP tools, and running batch
    campaigns.
  version: v1
servers:
  - url: https://api.vocily.ai
    description: Production
security: []
paths:
  /v1/speech:
    post:
      tags:
        - speech
      summary: Synthesize speech
      description: >-
        Turn text into audio with any voice from [`GET
        /v1/voices`](/developers/catalogues/voices), and get the audio back in
        the response body. Up to 1,000 characters per request.


        It is the same engine your agent speaks with, so it is the way to hear a
        voice before you put it on one.


        Nothing is stored. There is no id and no URL to fetch the audio again,
        so save the bytes if you need them. Each provider ships one model and it
        is chosen for you.
      operationId: create_speech_v1_speech_post
      parameters: []
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/PublicSpeechRequest'
            example:
              text: >-
                Hi, this is Riya calling from Acme about your appointment
                tomorrow.
              provider: cartesia
              voice_id: f6141af3-5f94-418c-80ed-a45d450e7e2e
              language: en
      responses:
        '200':
          description: The spoken audio — the response body is the audio file itself.
          content:
            audio/mpeg: {}
        '400':
          description: >-
            `provider` is not a voice provider; the voice or language was
            refused by it; or a control was sent to a provider that does not
            honour it (`speed` is not supported by `elevenlabs`, `emotion` only
            by `cartesia`). The message names the providers that do.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
              example:
                detail: >-
                  `provider` is not a voice provider; the voice or language was
                  refused by it; or a control was sent to a provider that does
                  not honour it (`speed` is not supported by `elevenlabs`,
                  `emotion` only by `cartesia`). The message names the providers
                  that do.
                code: BAD_REQUEST
        '401':
          description: Missing or invalid API key.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
              example:
                detail: Invalid API key
                code: UNAUTHORIZED
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
        '429':
          description: Rate limit exceeded — honor `Retry-After`.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
              example:
                code: rate_limited
        '502':
          description: The voice provider refused the request or returned no audio.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
              example:
                detail: The voice provider refused the request or returned no audio.
        '504':
          description: The voice provider did not respond in time.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
              example:
                detail: The voice provider did not respond in time.
      security:
        - bearerAuth: []
components:
  schemas:
    PublicSpeechRequest:
      properties:
        text:
          type: string
          maxLength: 1000
          minLength: 1
          title: Text
          description: What to say. Up to 1000 characters.
        provider:
          type: string
          title: Provider
          description: Voice provider, from `GET /v1/voices`.
        voice_id:
          type: string
          title: Voice Id
          description: >-
            Which voice, from `GET /v1/voices?provider=`. Must belong to
            `provider` — a voice id means nothing without the engine it came
            from.
        model:
          anyOf:
            - type: string
            - type: 'null'
          title: Model
          description: >-
            Read-only in practice. Each provider ships exactly one model and it
            is filled in from `provider`; `GET /v1/tts-capabilities` names
            which.
        language:
          anyOf:
            - type: string
            - type: 'null'
          title: Language
          description: >-
            The language to speak `text` in, e.g. `hi-IN`. Defaults to the
            provider's own default. A voice not certified for it is refused
            rather than substituted — `languages` on each `GET /v1/voices` row
            is that list.
        speed:
          anyOf:
            - type: number
              maximum: 4
              minimum: 0.25
            - type: 'null'
          title: Speed
          description: >-
            Speaking rate. 1.0 is the voice's natural pace, and each provider
            clamps it to its own range. **Not supported by `elevenlabs`** — sent
            for that provider it is refused rather than ignored.
        emotion:
          anyOf:
            - type: string
            - type: 'null'
          title: Emotion
          description: >-
            Emotional colour: `neutral` `curious` `excited` `enthusiastic`
            `happy` `content` `calm` `confident` `sad` `apologetic`
            `frustrated`. **Cartesia only** — sent for any other provider it is
            refused, because nothing else would speak it.
      additionalProperties: false
      type: object
      required:
        - text
        - provider
        - voice_id
      title: PublicSpeechRequest
      description: The body of `POST /v1/speech`.
    ApiError:
      type: object
      description: >-
        Error envelope. `code` is derived from the HTTP status, so branch on it
        for the CLASS of failure; the specific reason is `detail.code`. Every
        public refusal carries both.
      properties:
        detail:
          type: object
          description: >-
            The reason. `code` is the domain reason (e.g. `call_not_found`) and
            `message` is a sentence safe to log. On a `422` it also carries
            `errors[]`, one entry per rejected field — see
            `HTTPValidationError`.
          properties:
            code:
              type: string
              example: call_not_found
            message:
              type: string
              example: Call not found
          required:
            - code
            - message
        code:
          type: string
          description: Derived from the HTTP status, not the domain reason.
          example: NOT_FOUND
    HTTPValidationError:
      type: object
      title: HTTPValidationError
      description: >-
        A request the API could not read: a field of the wrong type, out of
        range, missing, or one we do not accept. Same envelope as every other
        error.
      properties:
        detail:
          type: object
          description: >-
            What was wrong, as `code`, a one-line `message`, and every offending
            field in `errors`.
          required:
            - code
            - message
            - errors
          properties:
            code:
              type: string
              enum:
                - validation_error
            message:
              type: string
              description: >-
                The first problem in one line, with a count of the rest — e.g.
                `model.temperature: Input should be less than or equal to 2 (and
                1 more)`.
            errors:
              type: array
              items:
                $ref: '#/components/schemas/ValidationError'
              description: >-
                One entry per offending field. **Every problem is reported at
                once**, not just the first, so a malformed body needs one round
                trip to fix rather than one per field.
        code:
          type: string
          enum:
            - VALIDATION_ERROR
          description: Derived from the HTTP status, as on every error.
    ValidationError:
      type: object
      title: ValidationError
      required:
        - field
        - message
        - type
      properties:
        field:
          type: string
          description: >-
            The offending field as a path from the root of your request —
            `voice.speed`, `variables[0].key`, or `query.limit` for a query
            parameter. **This is the field to read.**
        message:
          type: string
          description: What is wrong with it, in plain language.
        type:
          type: string
          description: >-
            A stable machine code for the kind of failure, e.g.
            `extra_forbidden` for a field we do not accept, `missing` for a
            required one, or `less_than_equal` for a number out of range. Switch
            on this rather than on `message`, which may be reworded.
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: 'Your API key as a Bearer token, e.g. `Authorization: Bearer vk_…`.'

````