> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server. # Video Translation API Guide > Build an end-to-end video translation pipeline with transcription, translation, text-to-speech, and AI lip sync using the sync. labs API. This guide shows a manually orchestrated video translation pipeline: transcribe the original audio, translate the text, generate speech in the target language, and use sync. labs to lip sync the new audio to the original video. Use this approach when you need control over the intermediate steps. > **Note** > > For translation and lipsync in a single API request, use [native Dubbing](/docs/tutorials/dubbing): send a video with its original audio embedded and set `dubParams` with `providerName` and `targetLang` on `POST /v2/generate`. Standard lipsync requests without `dubParams` match lip movements to the audio you supply; selecting a lipsync model does not translate that audio. If you already have translated audio, submit it as the audio input in a standard lipsync request instead. ## Full Pipeline Walkthrough #### Transcribe the original audio Extract the spoken words from your source video. OpenAI's Whisper is a solid choice for transcription. **`transcribe.py`** ```python transcribe.py from openai import OpenAI client = OpenAI() # Extract audio from video first (using ffmpeg or similar) audio_file = open("original-audio.wav", "rb") transcript = client.audio.transcriptions.create( model="whisper-1", file=audio_file, response_format="verbose_json", timestamp_granularities=["segment"], ) print(transcript.text) # Save segments with timestamps for alignment for segment in transcript.segments: print(f"[{segment['start']:.1f}s - {segment['end']:.1f}s] {segment['text']}") ``` The segment timestamps are useful for aligning translated audio with the correct video sections, especially for multi-speaker or long-form content. #### Translate the transcript Translate the transcribed text into the target language. Use a translation API or LLM for this step. **`translate.py`** ```python translate.py from openai import OpenAI client = OpenAI() original_text = "Welcome to our platform. Today we'll walk through the new features." target_language = "Spanish" response = client.chat.completions.create( model="gpt-4o", messages=[ { "role": "system", "content": f"Translate the following text to {target_language}. " f"Keep the tone natural and conversational. " f"Return only the translated text.", }, {"role": "user", "content": original_text}, ], ) translated_text = response.choices[0].message.content print(translated_text) # "Bienvenidos a nuestra plataforma. Hoy repasaremos las nuevas funciones." ``` > **Note** > > For production pipelines, consider specialized translation APIs (DeepL, Google Translate) for higher throughput and language coverage. #### Generate speech in the target language Convert the translated text to audio using a TTS service. ElevenLabs supports multilingual voice cloning -- you can clone the original speaker's voice and generate speech in the new language. **`generate_speech.py`** ```python generate_speech.py import requests ELEVENLABS_API_KEY = "your-elevenlabs-key" VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # Or a cloned voice ID response = requests.post( f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}", headers={ "xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json", }, json={ "text": "Bienvenidos a nuestra plataforma. Hoy repasaremos las nuevas funciones.", "model_id": "eleven_multilingual_v2", "voice_settings": { "stability": 0.5, "similarity_boost": 0.75, }, }, ) with open("translated-audio.mp3", "wb") as f: f.write(response.content) ``` Upload the generated audio to a publicly accessible URL for the next step. #### Lip sync with Sync Labs API Send the original video and the translated audio to Sync Labs. The API generates new lip movements matching the translated speech. **`lipsync.ts`** ```typescript lipsync.ts import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://your-cdn.com/original-video.mp4" }, { type: "audio", url: "https://your-cdn.com/translated-audio.mp3" }, ], model: "lipsync-2", options: { sync_mode: "cut_off" }, }); const jobId = response.id; console.log(`Lipsync job submitted: ${jobId}`); // Poll for completion let generation = await sync.generations.get(jobId); while (!["COMPLETED", "FAILED", "REJECTED"].includes(generation.status)) { await new Promise((r) => setTimeout(r, 10000)); generation = await sync.generations.get(jobId); } if (generation.status === "COMPLETED") { console.log(`Translated video ready: ${generation.outputUrl}`); } else { console.log(`Generation failed: ${jobId}`); } ``` **`lipsync.py`** ```python lipsync.py import time from sync import Sync from sync.common import Audio, Video, GenerationOptions sync = Sync() response = sync.generations.create( input=[ Video(url="https://your-cdn.com/original-video.mp4"), Audio(url="https://your-cdn.com/translated-audio.mp3"), ], model="lipsync-2", options=GenerationOptions(sync_mode="cut_off"), ) job_id = response.id print(f"Lipsync job submitted: {job_id}") # Poll for completion generation = sync.generations.get(job_id) while generation.status not in ["COMPLETED", "FAILED", "REJECTED"]: time.sleep(10) generation = sync.generations.get(job_id) if generation.status == "COMPLETED": print(f"Translated video ready: {generation.output_url}") else: print(f"Generation failed: {job_id}") ``` #### Use webhooks for production For production pipelines, replace polling with [webhooks](/api-reference/guides/webhooks). Pass a `webhookUrl` when creating the generation and Sync Labs sends a POST request when the job finishes. ```python response = sync.generations.create( input=[ Video(url="https://your-cdn.com/original-video.mp4"), Audio(url="https://your-cdn.com/translated-audio.mp3"), ], model="lipsync-2", webhook_url="https://your-app.com/webhooks/sync", ) ``` ## Shortcut: Built-in ElevenLabs Integration You can skip the separate TTS step by using Sync Labs' built-in ElevenLabs integration. Pass the translated text directly and Sync Labs handles TTS and lipsync in one call. ```python from sync import Sync from sync.common import Video, TTS, GenerationOptions sync = Sync() response = sync.generations.create( input=[ Video(url="https://your-cdn.com/original-video.mp4"), TTS( provider={ "name": "elevenlabs", "voiceId": "EXAVITQu4vr4xnSDxMaL", "script": "Bienvenidos a nuestra plataforma. Hoy repasaremos las nuevas funciones.", "stability": 0.5, "similarityBoost": 0.75, } ), ], model="lipsync-2", options=GenerationOptions(sync_mode="cut_off"), ) ``` See the [Integrations](/docs/product/integrations) page for setup instructions and voice configuration. ## Using the sync-examples Repository For a complete, ready-to-run translation pipeline, check the [sync-examples repository](https://github.com/synchronicity-labs/sync-examples). The translation example includes transcription with Whisper, translation with GPT, TTS with ElevenLabs, and lipsync with Sync Labs -- all wired together. ```bash git clone https://github.com/synchronicity-labs/sync-examples.git cd sync-examples/translation/python pip install -r requirements.txt # Configure your API keys in args.py python main.py ``` ## Quality Optimization Tips #### Choose the right model Use **[lipsync-2](/models/lipsync)** for standard translation jobs. Use **[lipsync-2-pro](/models/lipsync)** for premium content where facial detail (beards, teeth, wrinkles) matters. The quality difference is most visible in close-up shots. #### Ensure audio quality Clean, high-quality TTS audio produces better lipsync results. Use high-fidelity TTS models (like `eleven_multilingual_v2`) and avoid noisy or compressed audio files. #### Match speaking pace Translated text often has a different word count than the original. Tune your TTS speed settings so the translated audio duration roughly matches the original video length. This reduces artifacts from `sync_mode` adjustments. ## Handling Long Videos For videos longer than a few minutes, break them into segments: 1. **Transcribe with timestamps** -- Use Whisper's segment output to identify natural break points. 2. **Translate segment by segment** -- Translate each chunk individually for better accuracy. 3. **Generate audio per segment** -- Create separate TTS audio files for each segment. 4. **Use the Segments API** -- Submit all segments in a single Sync Labs API call with different audio inputs per time range. See the [Segments Guide](/developer-guides/segments). For batch translation of multiple videos, use the [Batch API](/api-reference/guides/batch-processing) to submit up to 500 generation requests in one operation. > Build an end-to-end video translation pipeline with transcription, translation, text-to-speech, and AI lip sync using the sync. labs API.