> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server.

# Text-to-Speech Lip Sync Guide

> Combine text-to-speech with AI lip sync to create talking head videos programmatically. Integration guide for ElevenLabs and other TTS providers with the sync. labs API.

Combine text-to-speech with lip sync to create talking head videos from just text and a source video. Type a script, pick a voice, and Sync Labs generates a video where the speaker's lips match the spoken words.

## Choose the right dubbing workflow

| Starting point                                          | Workflow                                                                                                                                                                                                                                                                         |
| :------------------------------------------------------ | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| A source video with speech that needs translation       | Use [`dubParams`](/tutorials/video-dubbing-api-guide#built-in-api-dubbing-with-dubparams) to translate, synthesize speech, and lip sync in one job. Do not add a separate audio or text input.                                                                                   |
| A translated script and a specific ElevenLabs voice ID  | Use a `video` input plus the `text` input below. For a voice in your own ElevenLabs account, connect that account in [Integrations settings](https://sync.so/settings/integrations); connection eligibility depends on your plan and the integration's requirements shown there. |
| A finished recording from a voice actor or TTS provider | Submit the source `video` and finished `audio` using the [external TTS workflow](#using-external-tts-providers).                                                                                                                                                                 |

`dubParams` uses the source video's audio for automatic dubbing. Use the text-input workflow when you need to select a specific ElevenLabs voice.

## Using the Built-in ElevenLabs Integration

The fastest path. Sync Labs' ElevenLabs integration handles TTS and lipsync in a single API call -- no need to generate and host audio separately.

**`tts_lipsync.ts`**

```typescript tts_lipsync.ts
import { SyncClient } from "@sync.so/sdk";

const sync = new SyncClient();

async function main() {
    const response = await sync.generations.create({
        input: [
            {
                type: "video",
                url: "https://assets.sync.so/docs/example-video.mp4",
            },
            {
                type: "text",
                provider: {
                    name: "elevenlabs",
                    voiceId: "EXAVITQu4vr4xnSDxMaL",
                    script: "Hey there. I wanted to walk you through our latest features. We shipped three major updates this week.",
                    stability: 0.5,
                    similarityBoost: 0.75,
                },
            },
        ],
        model: "lipsync-2",
        options: { sync_mode: "cut_off" },
    });

    const jobId = response.id;
    console.log(`Job submitted: ${jobId}`);

    // Poll for completion
    let generation = await sync.generations.get(jobId);
    while (!["COMPLETED", "FAILED", "REJECTED"].includes(generation.status)) {
        await new Promise((r) => setTimeout(r, 10000));
        generation = await sync.generations.get(jobId);
    }

    if (generation.status === "COMPLETED") {
        console.log(`Video ready: ${generation.outputUrl}`);
    } else {
        console.log(`Generation failed: ${jobId}`);
    }
}

main();
```

**`tts_lipsync.py`**

```python tts_lipsync.py
import time
from sync import Sync
from sync.common import Video, TTS, GenerationOptions

sync = Sync()

response = sync.generations.create(
    input=[
        Video(url="https://assets.sync.so/docs/example-video.mp4"),
        TTS(
            provider={
                "name": "elevenlabs",
                "voiceId": "EXAVITQu4vr4xnSDxMaL",
                "script": "Hey there. I wanted to walk you through our latest features. We shipped three major updates this week.",
                "stability": 0.5,
                "similarityBoost": 0.75,
            }
        ),
    ],
    model="lipsync-2",
    options=GenerationOptions(sync_mode="cut_off"),
)

job_id = response.id
print(f"Job submitted: {job_id}")

# Poll for completion
generation = sync.generations.get(job_id)
while generation.status not in ["COMPLETED", "FAILED", "REJECTED"]:
    time.sleep(10)
    generation = sync.generations.get(job_id)

if generation.status == "COMPLETED":
    print(f"Video ready: {generation.output_url}")
else:
    print(f"Generation failed: {job_id}")
```

### ElevenLabs Provider Parameters

| Parameter         | Type   | Default | Description                                                    |
| :---------------- | :----- | :-----: | :------------------------------------------------------------- |
| `name`            | string |    --   | Must be `"elevenlabs"`                                         |
| `voiceId`         | string |    --   | ElevenLabs voice ID. Must be a non-empty string.               |
| `script`          | string |    --   | Text to speak (max 5,000 characters)                           |
| `stability`       | float  |   0.5   | Voice stability (0.0-1.0). Lower = more expressive.            |
| `similarityBoost` | float  |   0.75  | Voice similarity to original (0.0-1.0). Higher = closer match. |

> **Note**
>
> Enable the ElevenLabs integration from your [Integrations settings](https://sync.so/settings/integrations). You can use Sync Labs' built-in integration or provide your own ElevenLabs API key (Creator plan or higher).

> **Note**
>
> Some ElevenLabs voices require verification before they can be used for TTS. If a generation fails with a message that the voice may violate ElevenLabs Terms of Service or requires verification, try a different verified voice or complete verification with ElevenLabs before retrying.

## Using External TTS Providers

If you use a TTS provider other than ElevenLabs -- Google Cloud TTS, Amazon Polly, Azure Speech, or any other service -- generate the audio first, host it at a public URL, then pass it to Sync Labs.

**`external_tts.ts`**

```typescript external_tts.ts
import { SyncClient } from "@sync.so/sdk";

const sync = new SyncClient();

// Audio generated by your TTS provider, hosted at a public URL
const ttsAudioUrl = "https://your-cdn.com/generated-speech.mp3";

const response = await sync.generations.create({
    input: [
        { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" },
        { type: "audio", url: ttsAudioUrl },
    ],
    model: "lipsync-2",
    options: { sync_mode: "cut_off" },
});

const jobId = response.id;
console.log(`Job submitted: ${jobId}`);

let generation = await sync.generations.get(jobId);
while (!["COMPLETED", "FAILED", "REJECTED"].includes(generation.status)) {
    await new Promise((r) => setTimeout(r, 10000));
    generation = await sync.generations.get(jobId);
}

if (generation.status === "COMPLETED") {
    console.log(`Video ready: ${generation.outputUrl}`);
}
```

**`external_tts.py`**

```python external_tts.py
import time
from sync import Sync
from sync.common import Audio, Video, GenerationOptions

sync = Sync()

# Audio generated by your TTS provider, hosted at a public URL
tts_audio_url = "https://your-cdn.com/generated-speech.mp3"

response = sync.generations.create(
    input=[
        Video(url="https://assets.sync.so/docs/example-video.mp4"),
        Audio(url=tts_audio_url),
    ],
    model="lipsync-2",
    options=GenerationOptions(sync_mode="cut_off"),
)

job_id = response.id
print(f"Job submitted: {job_id}")

# Poll for completion
generation = sync.generations.get(job_id)
while generation.status not in ["COMPLETED", "FAILED", "REJECTED"]:
    time.sleep(10)
    generation = sync.generations.get(job_id)

if generation.status == "COMPLETED":
    print(f"Video ready: {generation.output_url}")
```

This approach works with any TTS provider. The only requirement is that the audio file is accessible via a public URL.

## Voice Cloning Workflow

Clone a speaker's voice with ElevenLabs, then use that cloned voice ID with Sync Labs' integration. The result: a video where the speaker looks AND sounds like themselves -- speaking entirely new words.

#### Clone the voice

Upload a clean audio sample of the speaker to ElevenLabs to create a cloned voice. You get back a voice ID.

```python
# Use ElevenLabs API or dashboard to clone a voice
# https://elevenlabs.io/docs/voices/voice-cloning
cloned_voice_id = "your-cloned-voice-id"
```

#### Generate lipsync with the cloned voice

Use the cloned voice ID in your Sync Labs API call.

```python
from sync import Sync
from sync.common import Video, TTS, GenerationOptions

sync = Sync()

response = sync.generations.create(
    input=[
        Video(url="https://your-cdn.com/speaker-video.mp4"),
        TTS(
            provider={
                "name": "elevenlabs",
                "voiceId": cloned_voice_id,
                "script": "This is the new script I want the speaker to say.",
                "stability": 0.5,
                "similarityBoost": 0.85,  # Higher similarity for cloned voices
            }
        ),
    ],
    model="lipsync-2-pro",  # Pro model for highest quality
    options=GenerationOptions(sync_mode="cut_off"),
)
```

#### Download the result

Poll for completion and retrieve the output video. The speaker now says the new script with their own voice and matching lip movements.

## Best Practices

#### Keep scripts under 5,000 characters

The ElevenLabs integration has a 5,000-character limit per generation. For longer scripts, split them into segments using the [Segments API](/developer-guides/segments), with each segment referencing a separate TTS input.

#### Tune voice settings

**Stability** controls how consistent the voice sounds. Lower values (0.2-0.4) produce more expressive, varied speech. Higher values (0.6-0.8) produce more consistent, predictable speech. **Similarity boost** controls how closely the output matches the original voice. For cloned voices, use higher values (0.8-0.9).

#### Use sync-3 for the highest quality

For talking head videos where the face is prominent or the scene is challenging, [sync-3](/models/sync-3) is the highest-quality option. It handles obstructions, extreme angles, low light, and 4K output better than earlier lipsync models. Use [lipsync-2-pro](/models/lipsync) when you want premium facial detail at a lower price point than sync-3.

#### Use react-1 for expressive results

For short clips (under 15 seconds) where you want the speaker to show emotion, use [react-1](/models/react) with an emotion prompt. The model generates facial expressions and head movements that match the audio tone.

## Multi-Segment TTS

For longer scripts or multi-speaker scenarios, use the Segments API with multiple TTS inputs:

```python
from sync import Sync
from sync.common import Video, TTS

sync = Sync()

response = sync.generations.create(
    input=[
        Video(url="https://your-cdn.com/video.mp4"),
        TTS(
            provider={
                "name": "elevenlabs",
                "voiceId": "voice-id-1",
                "script": "Welcome to the first section of our presentation.",
            },
            ref_id="intro",
        ),
        TTS(
            provider={
                "name": "elevenlabs",
                "voiceId": "voice-id-2",
                "script": "Now let me hand it over to my colleague for the demo.",
            },
            ref_id="handoff",
        ),
    ],
    segments=[
        {"startTime": 0, "endTime": 8, "audioInput": {"refId": "intro"}},
        {"startTime": 8, "endTime": 15, "audioInput": {"refId": "handoff"}},
    ],
    model="lipsync-2",
)
```

## Troubleshooting TTS

#### Why did my TTS generation fail?

TTS generation failures typically come from one of four issues: an empty or invalid voice ID, a script that exceeds the character limit, or a missing ElevenLabs API key configuration. First, verify that the `voiceId` you are passing is a valid, non-empty ElevenLabs voice ID -- passing an empty string (`""`) returns a 422 error. Voice IDs can also expire if the voice is deleted from your ElevenLabs account or if you are referencing a shared voice that is no longer available. Second, check that your `script` field is under the 5,000-character limit; scripts that exceed this limit will be rejected. Third, confirm that the ElevenLabs integration is enabled in your [Integrations settings](https://sync.so/settings/integrations). Free accounts use Sync Labs' built-in ElevenLabs key, while Creator plan and above can provide their own API key for higher quotas. Check the error response from the [GET /v2/generate/\{id}](/api-reference/api/generate-api/get) endpoint for the specific error code and message.

#### How do I find my ElevenLabs voice ID?

To find your ElevenLabs voice ID, log in to the [ElevenLabs dashboard](https://elevenlabs.io) and navigate to the Voices section. Select the voice you want to use, then look for the voice ID in the URL bar or in the voice settings panel -- it is a string of characters like `EXAVITQu4vr4xnSDxMaL`. You can also find voice IDs through the ElevenLabs API by calling their [List Voices endpoint](https://elevenlabs.io/docs/api-reference/voices). If you are using a cloned voice, the voice ID is returned when you create the clone. Copy the voice ID exactly as shown and pass it as the `voiceId` parameter in your Sync Labs API request. Note that voice IDs are case-sensitive. If you are using Sync Labs' built-in ElevenLabs integration on a free account, you can use any of the default ElevenLabs voices without needing your own ElevenLabs account.

#### Can I use TTS in languages other than English?

Yes, Sync Labs' TTS integration supports multiple languages through ElevenLabs. ElevenLabs offers multilingual voice models that can generate speech in over 29 languages including Spanish, French, German, Portuguese, Japanese, Chinese, Arabic, Hindi, and many more. To use TTS in a non-English language, choose an ElevenLabs voice that supports your target language -- multilingual voices are labeled as such in the ElevenLabs voice library. Write your `script` in the target language and the TTS engine will generate speech in that language. The lip sync model will then match the lip movements to the generated audio regardless of the language, as Sync Labs' lip sync models are language-agnostic. For the best results, select a voice that is native to your target language rather than relying on a single voice to handle all languages.

## Text-to-Speech Lip Sync Reference

### TTS Provider Options

| Provider                                                      | Integration Type                                                         | Setup                                                                                                               | Plan Requirement                                                                |
| :------------------------------------------------------------ | :----------------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------ | :------------------------------------------------------------------------------ |
| ElevenLabs (built-in)                                         | Native -- single API call handles TTS and lip sync                       | Enable from Integrations settings at [https://sync.so/settings/integrations](https://sync.so/settings/integrations) | Free tier uses Sync Labs' built-in key; Creator+ can use own ElevenLabs API key |
| External (Google Cloud TTS, Amazon Polly, Azure Speech, etc.) | Two-step -- generate audio externally, then pass public URL to Sync Labs | Generate audio with your TTS provider, host at a public URL, pass URL as audio input                                | Any plan                                                                        |

### ElevenLabs Provider Parameters

| Parameter         | Type   | Default | Description                                                     |
| :---------------- | :----- | :-----: | :-------------------------------------------------------------- |
| `name`            | string |    --   | Must be `"elevenlabs"`                                          |
| `voiceId`         | string |    --   | ElevenLabs voice ID. Must be a non-empty string.                |
| `script`          | string |    --   | Text to speak (max 5,000 characters)                            |
| `stability`       | float  |   0.5   | Voice stability (0.0--1.0). Lower = more expressive.            |
| `similarityBoost` | float  |   0.75  | Voice similarity to original (0.0--1.0). Higher = closer match. |

### Key Limits

| Field                                  | Value                                                                                     |
| :------------------------------------- | :---------------------------------------------------------------------------------------- |
| Maximum script length                  | 5,000 characters per generation                                                           |
| Long script handling                   | Split into segments using the Segments API, each segment referencing a separate TTS input |
| Recommended model for TTS lip sync     | sync-3 for highest quality; lipsync-2-pro for premium detail at a lower price point       |
| Recommended model for expressive clips | react-1 (under 15 seconds, with emotion prompts)                                          |

### Voice Cloning Steps

| Step                   | Action                                                                                                                                                                                 |
| :--------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1. Clone the voice     | Upload a clean audio sample of the speaker to ElevenLabs to create a cloned voice. You receive a voice ID.                                                                             |
| 2. Generate lip sync   | Use the cloned voice ID in the Sync Labs API call with `provider.name: "elevenlabs"` and `provider.voiceId: "<cloned-voice-id>"`. Set `similarityBoost` to 0.8--0.9 for cloned voices. |
| 3. Download the result | Poll for completion and retrieve the output video. The speaker says the new script with their own voice and matching lip movements.                                                    |

### Common TTS Troubleshooting

| Problem                                   | Cause                                                            | Solution                                                                                                            |
| :---------------------------------------- | :--------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------ |
| Voice ID not found                        | Invalid or expired ElevenLabs voice ID                           | Verify the voice ID in your ElevenLabs dashboard; ensure the voice has not been deleted                             |
| voiceId must not be empty                 | Empty string passed as `voiceId` when provider is specified      | Provide a valid, non-empty ElevenLabs voice ID                                                                      |
| Input validation failed on TTS request    | Script exceeds 5,000-character limit or missing required fields  | Reduce script length to under 5,000 characters; ensure `name`, `voiceId`, and `script` are all provided             |
| Audio quality sounds robotic or unnatural | Stability set too high or similarityBoost set too low            | Lower stability to 0.2--0.4 for more expressive speech; raise similarityBoost to 0.8+ for cloned voices             |
| ElevenLabs integration not working        | Integration not enabled in account settings                      | Enable from Integrations settings at [https://sync.so/settings/integrations](https://sync.so/settings/integrations) |
| TTS generation fails silently             | Using external provider but audio URL is not publicly accessible | Ensure the audio file is hosted at a public URL with no authentication required                                     |
| Lip sync quality is poor with TTS audio   | Using a lower-quality model                                      | Switch to sync-3 for the highest quality, or lipsync-2-pro for premium detail at a lower price point                |