> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server.

# Text-to-Speech

> Synthesize speech from a script with sync. labs and get back a hosted audio URL you can lip sync onto any video.

`POST /v2/tts` synthesizes speech from a script and returns a hosted audio URL. Unlike the [built-in TTS lip sync flow](/tutorials/text-to-speech-lipsync), this endpoint is standalone: it does not run lip sync. You get back a stable `url` for the synthesized audio, which you can preview, store, or — most usefully — reuse as an `audio` input in [`POST /v2/generate`](/api-reference/api/generate-api/create) to lip sync that exact take onto a video.

For durable submission and polling, see [Asynchronous Text-to-Speech](/developer-guides/async-text-to-speech). The synchronous endpoint below keeps its `200` response shape.

## Request

Send a JSON body with the script and the voice to synthesize with.

**`script`** `string` — required

The text to synthesize into speech, from 1 to 5,000 characters.

---

**`voiceId`** `string` — required

A voice id to synthesize with — either an ElevenLabs voice id (discover via [`GET /v2/voices`](/api-reference/api/voices-api/list)) or the id of a voice cloned via [`POST /v2/voices`](/developer-guides/voice-cloning).

---

**`provider`** `string` — required

Use `elevenlabs` for ElevenLabs voices.

---

**`stability`** `double`

Voice stability (0–1). Higher is more consistent, lower is more expressive.

---

**`similarityBoost`** `double`

How closely the synthesized audio matches the original voice (0–1).

---

## Response

A `200` response returns the synthesized take.

**`id`** `string` — required

A unique identifier for the synthesized audio.

---

**`url`** `string` — required

The hosted URL of the synthesized audio. Reuse it as an `audio` input in `POST /v2/generate` to keep the same take across generations.

---

**`duration`** `double` — required

Duration of the synthesized audio in seconds.

---

## Synthesize speech

Lead with the script and a `voiceId`. The `curl` example below is the source of truth for the request shape.

**`curl`**

```bash curl
curl -X POST https://api.sync.so/v2/tts \
  -H "x-api-key: $SYNC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "script": "Hey there. I wanted to walk you through our latest features.",
    "provider": "elevenlabs",
    "voiceId": "EXAVITQu4vr4xnSDxMaL",
    "stability": 0.5,
    "similarityBoost": 0.75
  }'
```

**`tts.py`**

```python tts.py
from sync import Sync

sync = Sync()

result = sync.tts.create(
    script="Hey there. I wanted to walk you through our latest features.",
    provider="elevenlabs",
    voice_id="EXAVITQu4vr4xnSDxMaL",
    stability=0.5,
    similarity_boost=0.75,
)

print(f"Synthesized audio: {result.url} ({result.duration}s)")
```

**`tts.ts`**

```typescript tts.ts
import { SyncClient } from "@sync.so/sdk";

const sync = new SyncClient();

const result = await sync.tts.create({
    script: "Hey there. I wanted to walk you through our latest features.",
    provider: "elevenlabs",
    voiceId: "EXAVITQu4vr4xnSDxMaL",
    stability: 0.5,
    similarityBoost: 0.75,
});

console.log(`Synthesized audio: ${result.url} (${result.duration}s)`);
```

A successful response looks like this:

```json
{
  "id": "6533643b-aceb-4c40-967e-d9ba9baac39e",
  "url": "https://assets.sync.so/docs/example-tts.mp3",
  "duration": 2.1
}
```

> **Note**
>
> `voiceId` accepts any ElevenLabs voice id — list the voices available to your organization with [`GET /v2/voices`](/api-reference/api/voices-api/list) — or the id of a voice you cloned via [`POST /v2/voices`](/developer-guides/voice-cloning). Voice ids are case-sensitive.

## Synthesize, then lip sync

The standalone endpoint is most powerful as the first half of a two-step flow: synthesize a take with `/v2/tts`, then pass the returned `url` as an `audio` input to `/v2/generate` to lip sync it onto a video.

#### Synthesize the take

Call `POST /v2/tts` and keep the returned `url`. This is your hosted audio take.

#### Lip sync it onto a video

Pass that `url` as an `audio` input in `POST /v2/generate`, alongside your video input.

#### Poll for completion

Poll `GET /v2/generate/{id}` until `status` is `COMPLETED`; `outputUrl` contains the lipsynced video.

**`curl`**

```bash curl
# 1. Synthesize the take and capture the hosted url
TTS_URL=$(curl -s -X POST https://api.sync.so/v2/tts \
  -H "x-api-key: $SYNC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "script": "Hey there. I wanted to walk you through our latest features.",
    "provider": "elevenlabs",
    "voiceId": "EXAVITQu4vr4xnSDxMaL"
  }' | jq -r '.url')

# 2. Lip sync the synthesized take onto a video
curl -X POST https://api.sync.so/v2/generate \
  -H "x-api-key: $SYNC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lipsync-2",
    "input": [
      { "type": "video", "url": "https://assets.sync.so/docs/example-video.mp4" },
      { "type": "audio", "url": "'"$TTS_URL"'" }
    ],
    "options": { "sync_mode": "cut_off" }
  }'
```

**`tts_then_generate.py`**

```python tts_then_generate.py
import time
from sync import Sync
from sync.common import Audio, Video, GenerationOptions

sync = Sync()

# 1. Synthesize the take and keep the hosted url
take = sync.tts.create(
    script="Hey there. I wanted to walk you through our latest features.",
    provider="elevenlabs",
    voice_id="EXAVITQu4vr4xnSDxMaL",
)

# 2. Lip sync the synthesized take onto a video
response = sync.generations.create(
    model="lipsync-2",
    input=[
        Video(url="https://assets.sync.so/docs/example-video.mp4"),
        Audio(url=take.url),
    ],
    options=GenerationOptions(sync_mode="cut_off"),
)

job_id = response.id
generation = sync.generations.get(job_id)
while generation.status not in ["COMPLETED", "FAILED", "REJECTED"]:
    time.sleep(10)
    generation = sync.generations.get(job_id)

if generation.status == "COMPLETED":
    print(f"Video ready: {generation.output_url}")
```

**`tts_then_generate.ts`**

```typescript tts_then_generate.ts
import { SyncClient } from "@sync.so/sdk";

const sync = new SyncClient();

// 1. Synthesize the take and keep the hosted url
const take = await sync.tts.create({
    script: "Hey there. I wanted to walk you through our latest features.",
    provider: "elevenlabs",
    voiceId: "EXAVITQu4vr4xnSDxMaL",
});

// 2. Lip sync the synthesized take onto a video
const response = await sync.generations.create({
    model: "lipsync-2",
    input: [
        { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" },
        { type: "audio", url: take.url },
    ],
    options: { sync_mode: "cut_off" },
});

let generation = await sync.generations.get(response.id);
while (!["COMPLETED", "FAILED", "REJECTED"].includes(generation.status)) {
    await new Promise((r) => setTimeout(r, 10000));
    generation = await sync.generations.get(response.id);
}

if (generation.status === "COMPLETED") {
    console.log(`Video ready: ${generation.outputUrl}`);
}
```

> **Note**
>
> The generation response echoes the synthesized take as `synthesizedAudioUrl`, so you can reuse the exact same audio across multiple generations without re-synthesizing. This is also present when you submit a `text` input directly to `/v2/generate` — see the [built-in TTS lip sync flow](/tutorials/text-to-speech-lipsync).

## Quotas and rate limits

> **Warning**
>
> Free-tier API keys share a monthly ElevenLabs allowance of **10 synthesis operations across TTS and dubbing combined**. Paid plans are billed per use. See the [Billing](/docs/product/billing) page for details.

`POST /v2/tts` is rate limited to **60 requests per minute per key**. Exceeding the limit returns a `429` — back off and retry. See [Rate Limiting](/api-reference/guides/rate-limits) for the recommended retry strategy.

## FAQ

#### What's the difference between /v2/tts and TTS lip sync?

`POST /v2/tts` only synthesizes audio — it returns a hosted `url` and does not touch video. The [built-in TTS lip sync flow](/tutorials/text-to-speech-lipsync) passes a `text` input to `POST /v2/generate`, which synthesizes the audio and runs lip sync in a single call. Use `/v2/tts` when you want to inspect, reuse, or store the synthesized take before (or independent of) lip syncing it.

#### How do I find a voiceId?

Call [`GET /v2/voices`](/api-reference/api/voices-api/list) to list the voices available to your organization, including built-in ElevenLabs voices and any you have cloned. Each entry includes an `id` you can pass as `voiceId`. To create your own voice, see [Voice Cloning](/developer-guides/voice-cloning).

#### Can I reuse the same take across multiple videos?

Yes. The `url` returned by `/v2/tts` is stable — pass it as an `audio` input to as many `POST /v2/generate` calls as you like. The generation response also echoes it as `synthesizedAudioUrl`, so you can recover the exact take from a completed generation without re-synthesizing.

#### How do I handle longer scripts or multiple speakers?

Synthesize each section as a separate `/v2/tts` take, then assign the resulting audio urls to different time ranges with the [Segments API](/developer-guides/segments). Each segment can reference a different audio input, which is how you build multi-speaker and long-form lip sync in a single generation.

## Related

* [Voice Cloning](/developer-guides/voice-cloning) — clone a custom voice and synthesize with its id.
* [Segments](/developer-guides/segments) — assign different synthesized takes to different parts of the timeline.
* [Text-to-Speech Lip Sync Guide](/tutorials/text-to-speech-lipsync) — synthesize and lip sync in a single `/v2/generate` call.