> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server. # Segments Guide > Multi-segment lip sync guide. sync. labs different audio clips to different parts of your video timeline in a single API call. ## Overview Segments let you sync different audio clips to different time ranges within a single video in one API call. This enables multi-speaker lip sync by letting you assign different audio inputs to different parts of your video. Using segments, you can: * LipSync different audio clips to different parts of your video * Use specific portion of audio input to lipsync a segment for precise timing * Use both audio and text-to-speech inputs to lipsync multiple segments with different input types in a single generation ## Basic Concepts To use segments feature, you need to provide a top-level [`segments`](/api-reference/api/generate-api/create#request.body.segments) array with each item defining a video time range/segment, each with its own audio configuration. ### Segment Each segment item takes the following properties: **`startTime`** `double` — required Segment start time in seconds --- **`endTime`** `double` — required Segment end time in seconds --- **`audioInput`** `SegmentAudioInput` — required Audio configuration with refId and optional cropping --- **`optionsOverride`** `SegmentOptionsOverride` Override generation options for this specific segment --- ### audioInput Each segment requires exactly one audioInput. audioInput takes the following properties: **`refId`** `string` — required Reference ID of the audio/text-to-speech input to use for this segment --- **`startTime`** `double` Optional start time (in seconds) to crop the referenced audio. When specified, endTime must also be provided --- **`endTime`** `double` Optional end time (in seconds) to crop the referenced audio. When specified, startTime must also be provided --- The specified audioInput will be used to lipsync the video segment between startTime and endTime. ### optionsOverride Each segment can optionally override the top-level generation options. This allows you to apply different settings per segment, including targeting different speakers. For multi-speaker videos, use `active_speaker_detection` to target a different person in each segment. See [Speaker Selection — API](/developer-guides/speaker-selection) for full details on speaker selection options. **`sync_mode`** `SyncMode` Override the [sync mode](/developer-guides/sync-mode) for this segment --- **`temperature`** `double` Override expressiveness (0-1) for this segment --- **`occlusion_detection_enabled`** `boolean` Override occlusion detection for this segment --- **`active_speaker_detection`** `ActiveSpeaker` Override [active speaker detection](/developer-guides/speaker-selection) for this segment. Useful when different segments have different speakers. Accepts the same options as the top-level `active_speaker_detection`: * `auto_detect`: automatically detect and target the active speaker * `v3`: use ASD v3 * `frame_number` + `coordinates`: manually specify the speaker by frame and point * `bounding_boxes`: provide per-frame bounding boxes if you have detection data * `bounding_boxes_url`: point to an external JSON file containing bounding boxes (recommended for long videos to avoid large request payloads) --- ## How segment timing & duration works When a segment's audio and its video window (`endTime` − `startTime`) are different lengths, [`sync_mode`](/compatibility-and-tips/media-content-tips#sync-mode-options) decides how the mismatch is resolved — and the mode you choose changes the **effective duration** of that segment in the output: | `sync_mode` | audio longer than the window | audio shorter than the window | | ----------- | ------------------------------------------------------- | -------------------------------------------- | | `bounce` | video bounces (forward then reverse) to cover the audio | window trimmed to the audio | | `loop` | video loops to cover the audio | window trimmed to the audio | | `cut_off` | output clipped to the **shorter** track (the window) | window trimmed to the audio | | `silence` | plays the full audio (video padded to cover it) | full window plays; audio padded with silence | | `remap` | video time-stretched to match the audio | window trimmed to the audio | In short, a segment's effective length is `cut_off` → min(audio, window), `silence` → max(audio, window), and `loop` / `bounce` / `remap` → the audio length. > **Note** > > The same rules apply to a single-segment (non-`segments`) generation, where the "window" is the whole video. A 30s audio on a 15s video with `sync_mode: cut_off` produces a **15s** output (clipped to the video); `silence` / `loop` / `bounce` / `remap` produce a \~30s output. ### Output can be longer than your source video Each segment places its audio take into its own window on the timeline, and **each window expands to fit its audio** (per the `sync_mode` above). When there are gaps between segments, those expanded windows shift the rest of the timeline — so the **output can be longer than the source video**. For example, two segments `[1s–3s]` and `[5s–8s]` on a 10s video can produce a \~16.8s output: each window stretched to fit its full audio take. This is expected behavior, not a bug. **To keep the output the same length as the source video**, crop each audio take to its window with `audioInput.startTime`/`endTime`, so the take is exactly as long as the window and the window doesn't expand: ```json { "startTime": 1, "endTime": 3, "audioInput": { "refId": "audio_1", "startTime": 0, "endTime": 2 } } ``` Here the 2-second audio slice exactly fills the 2-second window, so the timeline — and the output duration — stays aligned with the source. ## API Usage Examples #### Single Segment with Single Audio ```python from sync import Sync from sync.common import Audio, Video sync = Sync() response = sync.generations.create( input=[ Video(url="https://assets.sync.so/docs/example-video.mp4"), Audio(url="https://assets.sync.so/docs/example-audio.wav", ref_id="audio_1"), ], model="lipsync-2", segments=[ GenerationSegment( start_time=2, end_time=5, audio_input=SegmentAudioInput(ref_id="audio_1"), ), ], ) ``` ```typescript import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" }, { type: "audio", url: "https://assets.sync.so/docs/example-audio.wav", refId: "audio_1" }, ], segments: [ { startTime: 2, endTime: 5, audioInput: { refId: "audio_1" } }, ], model: "lipsync-2" }); ``` #### Multiple Segments with Single Audio ### Multiple Segments with Single Audio Input ```python from sync import Sync from sync.common import Audio, Video, TTS sync = Sync() response = sync.generations.create( input=[ Video(url="https://assets.sync.so/docs/example-video.mp4"), Audio(url="https://assets.sync.so/docs/example-audio.wav", ref_id="audio_1") ], segments=[ { "startTime": 2, "endTime": 5, "audioInput": {"refId": "audio_1", "startTime": 2, "endTime": 5} }, { "startTime": 6, "endTime": 8, "audioInput": {"refId": "audio_1", "startTime": 6, "endTime": 8} } ], model="lipsync-2" ) ``` ```typescript import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" }, { type: "audio", url: "https://assets.sync.so/docs/example-audio.wav", refId: "audio_1" }, ], segments: [ { startTime: 2, endTime: 5, audioInput: { refId: "audio_1", startTime: 2, endTime: 5 } }, { startTime: 6, endTime: 8, audioInput: { refId: "audio_1", startTime: 6, endTime: 8 } } ], model: "lipsync-2" }); ``` #### Multiple Segments with Multiple Audio ### Multiple Segments with Single Audio Input ```python from sync import Sync from sync.common import Audio, Video, TTS sync = Sync() response = sync.generations.create( input=[ Video(url="https://assets.sync.so/docs/example-video.mp4"), Audio(url="https://assets.sync.so/docs/example-audio.wav", ref_id="audio_1"), Audio(url="https://assets.sync.so/docs/example-audio.wav", ref_id="audio_2") ], segments=[ { "startTime": 2, "endTime": 5, "audioInput": {"refId": "audio_1", "startTime": 2, "endTime": 5} }, { "startTime": 6, "endTime": 8, "audioInput": {"refId": "audio_2", "startTime": 6, "endTime": 8} } ], model="lipsync-2" ) ``` ```typescript import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" }, { type: "audio", url: "https://assets.sync.so/docs/example-audio.wav", refId: "audio_1" }, { type: "audio", url: "https://assets.sync.so/docs/example-audio.wav", refId: "audio_2" }, ], segments: [ { startTime: 2, endTime: 5, audioInput: { refId: "audio_1", startTime: 2, endTime: 5 } }, { startTime: 6, endTime: 8, audioInput: { refId: "audio_2", startTime: 6, endTime: 8 } } ], model: "lipsync-2" }); ``` #### Segments with Options Override ### Segments with Per-Segment Options Use `optionsOverride` to apply different generation settings to each segment. ```python from sync import Sync from sync.common import Audio, Video sync = Sync() response = sync.generations.create( input=[ Video(url="https://assets.sync.so/docs/example-video.mp4"), Audio(url="https://assets.sync.so/docs/audio1.wav", ref_id="audio_1"), Audio(url="https://assets.sync.so/docs/audio2.wav", ref_id="audio_2") ], segments=[ { "startTime": 0, "endTime": 5, "audioInput": {"refId": "audio_1"}, "optionsOverride": { "sync_mode": "loop", "temperature": 0.3 } }, { "startTime": 5, "endTime": 10, "audioInput": {"refId": "audio_2"}, "optionsOverride": { "sync_mode": "cut_off", "temperature": 0.7 } } ], model="lipsync-2", options={"sync_mode": "bounce"} # Default for segments without override ) ``` ```typescript import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://assets.sync.so/docs/example-video.mp4" }, { type: "audio", url: "https://assets.sync.so/docs/audio1.wav", refId: "audio_1" }, { type: "audio", url: "https://assets.sync.so/docs/audio2.wav", refId: "audio_2" } ], segments: [ { startTime: 0, endTime: 5, audioInput: { refId: "audio_1" }, optionsOverride: { sync_mode: "loop", temperature: 0.3 } }, { startTime: 5, endTime: 10, audioInput: { refId: "audio_2" }, optionsOverride: { sync_mode: "cut_off", temperature: 0.7 } } ], model: "lipsync-2", options: { sync_mode: "bounce" } // Default for segments without override }); ``` #### Multi-Speaker Segments with Active Speaker Detection ### Target Different Speakers Per Segment Use `active_speaker_detection` in `optionsOverride` to target different speakers in each segment. This is helpful when a video has multiple people and different segments should lipsync to different speakers. ```python from sync import Sync from sync.common import Audio, Video sync = Sync() response = sync.generations.create( input=[ Video(url="https://assets.sync.so/docs/two-speakers.mp4"), Audio(url="https://assets.sync.so/docs/speaker-a.wav", ref_id="audio_a"), Audio(url="https://assets.sync.so/docs/speaker-b.wav", ref_id="audio_b") ], segments=[ { "startTime": 0, "endTime": 5, "audioInput": {"refId": "audio_a"}, "optionsOverride": { "active_speaker_detection": { "frame_number": 0, "coordinates": [200, 300] # Point on speaker A's face } } }, { "startTime": 5, "endTime": 10, "audioInput": {"refId": "audio_b"}, "optionsOverride": { "active_speaker_detection": { "frame_number": 150, "coordinates": [600, 300] # Point on speaker B's face } } } ], model="lipsync-2" ) ``` ```typescript import { SyncClient } from "@sync.so/sdk"; const sync = new SyncClient(); const response = await sync.generations.create({ input: [ { type: "video", url: "https://assets.sync.so/docs/two-speakers.mp4" }, { type: "audio", url: "https://assets.sync.so/docs/speaker-a.wav", refId: "audio_a" }, { type: "audio", url: "https://assets.sync.so/docs/speaker-b.wav", refId: "audio_b" } ], segments: [ { startTime: 0, endTime: 5, audioInput: { refId: "audio_a" }, optionsOverride: { activeSpeakerDetection: { frameNumber: 0, coordinates: [200, 300] // Point on speaker A's face } } }, { startTime: 5, endTime: 10, audioInput: { refId: "audio_b" }, optionsOverride: { activeSpeakerDetection: { frameNumber: 150, coordinates: [600, 300] // Point on speaker B's face } } } ], model: "lipsync-2" }); ``` See [Speaker Selection — API](/developer-guides/speaker-selection) for details on `active_speaker_detection` options. ## Best Practices ### Planning Your Segments 1. **Map your timeline**: Identify video segments and corresponding audio needs 2. **Prepare audio files**: Ensure audio quality and appropriate duration 3. **Test segment boundaries**: Verify smooth transitions between segments ### Audio Preparation * Use consistent audio quality across all segments and the video's audio. * For best results, ensure proper timing alignment with video segments. If segment duration and corresponding audio duration don't match, rely on [sync\_mode](/developer-guides/sync-mode) to determine how to handle the mismatch. ## Troubleshooting ### Common Errors #### "Multiple audio inputs are only allowed when using multi-segments" Provide a top-level `segments` array when using multiple audio or text inputs. ```python # ❌ This will fail response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio1.wav"), # Multiple audio without segments Audio(url="audio2.wav") ], model="lipsync-2" ) # ✅ This will work response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio1.wav", ref_id="a1"), Audio(url="audio2.wav", ref_id="a2") ], segments=[ {"start_time": 0, "end_time": 10, "audio_input": {"refId": "a1"}}, {"start_time": 10, "end_time": 20, "audio_input": {"refId": "a2"}} ], model="lipsync-2" ) ``` #### "Unable to resolve audio input URL" Ensure all audio inputs have valid `url` or `assetId` values and that referenced `refId` values exist in your audio or text inputs. ```python # ❌ Missing refId reference response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") ], segments=[ {"start_time": 0, "end_time": 10, "audio_input": {"refId": "missing"}} # Wrong refId ], model="lipsync-2" ) # ✅ Correct refId reference response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") ], segments=[ {"start_time": 0, "end_time": 10, "audio_input": {"refId": "audio1"}} # Correct refId ], model="lipsync-2" ) ``` #### "Segment at index X is missing a valid audioInput.refId" This error occurs when a segment's `audio_input` is missing a `refId` or the `refId` is empty. Each segment must reference a valid audio or text input through its `refId`. ```python # ❌ Missing refId in segment response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") ], segments=[ { "start_time": 0, "end_time": 10, "audio_input": {} # Missing refId } ], model="lipsync-2" ) # ✅ Include refId in segment response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") ], segments=[ { "start_time": 0, "end_time": 10, "audio_input": {"refId": "audio1"} # Valid refId } ], model="lipsync-2" ) ``` #### "Segment at index X references unknown refId" This error occurs when a segment references a `refId` that doesn't exist in your audio or text inputs. Ensure all referenced `refId` values match exactly with those defined in your inputs. ```python # ❌ Segment references unknown refId response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") # refId is "audio1" ], segments=[ { "start_time": 0, "end_time": 10, "audio_input": {"refId": "nonexistent"} # References unknown refId } ], model="lipsync-2" ) # ✅ Segment references existing refId response = sync.generations.create( input=[ Video(url="video.mp4"), Audio(url="audio.wav", ref_id="audio1") # refId is "audio1" ], segments=[ { "start_time": 0, "end_time": 10, "audio_input": {"refId": "audio1"} # References existing refId } ], model="lipsync-2" ) ``` #### "Invalid segment time range: startTime must be \<= endTime" Each segment's `startTime` must be less than or equal to its `endTime`. Zero-length segments (where `startTime` equals `endTime`) are allowed for use cases like zero-duration crop points. ```python # ❌ Invalid: startTime greater than endTime segments=[ { "startTime": 10, "endTime": 5, # endTime must be >= startTime "audioInput": {"refId": "audio1"} } ] # ✅ Valid: startTime less than endTime segments=[ { "startTime": 0, "endTime": 10, "audioInput": {"refId": "audio1"} } ] # ✅ Valid: startTime equals endTime (zero-length segment) segments=[ { "startTime": 5, "endTime": 5, "audioInput": {"refId": "audio1"} } ] ``` #### "Invalid segment frame range: startFrame must be \< endFrame" When specifying segment boundaries using frames instead of seconds, `startFrame` must be strictly less than `endFrame`. Unlike time-based segments which allow equal start and end times, frame-based segments require at least one frame of difference. ```python # ❌ Invalid: startFrame greater than or equal to endFrame segments=[ { "startFrame": 100, "endFrame": 0, # endFrame must be > startFrame "audioInput": {"refId": "audio1"} } ] # ❌ Invalid: startFrame equals endFrame (not allowed for frames) segments=[ { "startFrame": 50, "endFrame": 50, # endFrame must be > startFrame "audioInput": {"refId": "audio1"} } ] # ✅ Valid: startFrame less than endFrame segments=[ { "startFrame": 0, "endFrame": 100, "audioInput": {"refId": "audio1"} } ] ``` #### "Invalid audioInput crop range" When cropping audio within a segment, both `startTime` and `endTime` must be provided, and `startTime` must be less than or equal to `endTime`. ```python # ❌ Incomplete crop range segments=[ { "startTime": 0, "endTime": 10, "audioInput": { "refId": "audio1", "startTime": 5 # Missing endTime } } ] # ❌ Invalid: startTime greater than endTime segments=[ { "startTime": 0, "endTime": 10, "audioInput": { "refId": "audio1", "startTime": 15, "endTime": 5 # endTime must be >= startTime } } ] # ✅ Valid: complete crop range with startTime <= endTime segments=[ { "startTime": 0, "endTime": 10, "audioInput": { "refId": "audio1", "startTime": 5, "endTime": 15 } } ] # ✅ Valid: zero-duration crop point (startTime equals endTime) segments=[ { "startTime": 0, "endTime": 10, "audioInput": { "refId": "audio1", "startTime": 5, "endTime": 5 } } ] ``` #### "When using multi-segments, please provide at least one audio or text input" Ensure you have at least one audio input or text input with a valid `refId` when using segments. ```python # ❌ This will fail - no audio or text inputs response = sync.generations.create( input=[ Video(url="https://example.com/video.mp4") ], segments=[ {"start_time": 0, "end_time": 10, "audio_input": {"refId": "missing"}} ], model="lipsync-2" ) # ✅ This will work - includes text input response = sync.generations.create( input=[ Video(url="https://example.com/video.mp4"), TTS( provider={ "name": "elevenlabs", "voiceId": "EXAVITQu4vr4xnSDxMaL", "script": "Hello world" }, ref_id="text1" ) ], segments=[ {"start_time": 0, "end_time": 10, "audio_input": {"refId": "text1"}} ], model="lipsync-2" ) ``` > Multi-segment lip sync guide. sync. labs different audio clips to different parts of your video timeline in a single API call.