> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server.

# Improving Lip Sync Quality

> Tips and best practices for getting the best lip sync results from sync. labs' AI models. Learn how to optimize your input video and audio for natural-looking lip movements.

Lip sync quality depends on the quality of your input video, the clarity of your audio, the model you choose, and how you configure sync mode settings. Following these guidelines will help you get the most natural-looking results from every generation.

## Video Input Best Practices

The quality and composition of your input video has the biggest impact on lip sync output. Follow these guidelines for the best results.

#### Use a front-facing camera angle

Frontal or near-frontal face angles produce the best lip sync results on lipsync-2 and lipsync-2-pro. Extreme profile (side-view) shots make face detection unreliable and can cause distorted output on these models. sync-3 natively supports extreme face angles including profiles and over-the-shoulder shots, but front-facing angles still produce the best results across all models.

#### Ensure good, even lighting

Well-lit faces with even lighting produce the cleanest results. Avoid harsh shadows across the face, backlighting that silhouettes the speaker, or flickering light sources that change frame to frame.

#### Minimize face obstructions

Keep the speaker's mouth and lower face area clear of obstructions. Hands, microphones, hair, sunglasses, and other objects covering the mouth region degrade lip sync accuracy on lipsync-2 and lipsync-2-pro — if obstructions are unavoidable on these models, enable the `occlusion_detection_enabled` option. [sync-3](/models/sync-3) detects and handles obstructions automatically with no configuration needed.

#### Use stable footage

Shaky or jittery footage makes face tracking less reliable. Use a tripod or stabilized camera when possible. If using handheld footage, apply video stabilization in post-production before submitting to Sync Labs.

#### Keep one face in frame

For the best results, have a single speaker visible in the frame. If your video has multiple faces, use the [Speaker Selection](/developer-guides/speaker-selection) feature to target the correct person. Without speaker selection, the model may sync to the wrong face.

#### Meet resolution requirements

Use at least 480p resolution for reliable face detection. Higher resolutions up to 4K (4096x2160) are supported and can improve output quality. We recommend 1080p as the best balance of quality and processing speed. Use MP4 with H.264 codec for optimal compatibility.

#### Preserve color metadata

If you plan to composite the generated output back onto the source footage, export SDR BT.709 with explicit color tags and use H.264 `yuv444p` when color consistency matters. See [Preserving Color During Generation](/compatibility-and-tips/media-formats-support#preserving-color-during-generation) for details.

## Audio Input Best Practices

Clear, well-recorded audio is essential for accurate lip-to-speech alignment.

#### Use clean speech without background noise

Background music, crowd noise, and overlapping conversations degrade lip sync accuracy. Isolate the speaker's voice as much as possible. If your audio has background noise, use a noise reduction tool before submitting. For song lip sync, isolate and upload just the vocals track.

#### One speaker per audio track

Each audio input should contain a single speaker. For multi-speaker scenarios, use the [Segments API](/developer-guides/segments) to assign different audio tracks to different time ranges, each with its own speaker.

#### Match audio and video duration

For the most predictable results, keep audio and video durations close to each other. When durations differ, use the [`sync_mode`](/developer-guides/sync-mode) parameter to control how the mismatch is handled. `cut_off` uses the shorter input duration, `bounce` or `loop` can extend short video over longer audio, `silence` preserves long video when audio is shorter, and `remap` adjusts video speed.

**Recommended audio formats:** WAV or MP3. All major audio formats are supported -- see [Media Formats Support](/compatibility-and-tips/media-formats-support) for the full list. Sync Labs supports any language for lip sync.

## Model Selection for Quality

Each model has different strengths. Choose the right model based on your quality and speed priorities.

| Priority       | Recommended Model                              | Why                                                                                                               |
| :------------- | :--------------------------------------------- | :---------------------------------------------------------------------------------------------------------------- |
| Most powerful  | [sync-3](/models/lipsync)                      | 4K native output, built-in obstruction detection, extreme angle support, full-shot processing ($0.107-$0.133/sec) |
| Premium detail | [lipsync-2-pro](/models/lipsync#lipsync-2-pro) | Enhanced detail for beards, teeth, and facial features using diffusion-based super resolution ($0.067-$0.083/sec) |
| Best balance   | [lipsync-2](/models/lipsync)                   | Natural lip movements that preserve the speaker's unique speaking style ($0.04-$0.05/sec)                         |
| Expressions    | [react-1](/models/react)                       | Adds emotion, facial expressions, and head movement to match audio tone (max 15s clips)                           |

> **Note**
>
> For most use cases, **lipsync-2** provides the best balance of quality and speed. Use **lipsync-2-pro** when you need enhanced detail for beards, teeth, or fine facial features. Use **sync-3** for production-grade results — it handles close-ups, profile shots, obstructions, and complex scenes that other models struggle with.

## Common Quality Issues and Fixes

#### Lip movements don't match the audio

Audio-video misalignment typically stems from one of three causes: sync mode configuration, audio quality, or duration mismatch. First, check your `sync_mode` setting -- `cut_off` mode trims audio that extends beyond the video length, which works well for most cases. If audio and video durations are significantly different, the `remap` mode adjusts video playback speed to match, but large speed changes can look unnatural. Try `cut_off` if you are seeing drift. Second, ensure your audio is clean -- background music, overlapping speakers, and heavy noise make it harder for the model to align lip movements to the correct speech patterns. Use a noise reduction tool on your audio before submitting. Third, verify that the audio language matches what the model expects; Sync Labs supports all languages, but audio with mixed languages in a single track can cause alignment issues.

#### Quality degrades in long videos

Long videos are internally divided into 30-40 second chunks for processing. If chunk boundaries fall at points where the face is partially visible, moving rapidly, or absent, the output quality can degrade at those transitions. For the best results with long-form content, consider splitting your video into segments under 2 minutes using the [Segments API](/developer-guides/segments) and processing each segment separately. Use [lipsync-2-pro](/models/lipsync#lipsync-2-pro) for long-form content where quality is critical -- its diffusion-based super resolution handles transitions between chunks more gracefully. Also ensure the speaker's face is consistently visible and well-lit throughout the entire video. Rapid scene changes within chunks are a common cause of processing timeouts and quality drops.

#### Face looks distorted or has artifacts

Visual artifacts around the face region usually result from challenging input conditions. Start by ensuring the speaker's face has good, even lighting without harsh shadows -- uneven lighting causes the model to produce inconsistent skin tones across frames. Verify the face is front-facing or near-frontal, as extreme angles produce distortion on lipsync-2 and lipsync-2-pro. Remove any obstructions covering the mouth area, including hands, microphones, or hair. On lipsync-2 and lipsync-2-pro, enable `occlusion_detection_enabled` for better handling of partially hidden faces. [sync-3](/models/sync-3) handles obstructions and extreme angles automatically. For persistent artifacts, try [sync-3](/models/sync-3) for the most robust results, or [lipsync-2-pro](/models/lipsync#lipsync-2-pro) for diffusion-based super resolution around teeth, beards, and fine facial features. Check that your input resolution is at least 480p; very low-resolution faces make detection and generation less reliable.

#### Multiple faces but wrong one is synced

When a video contains multiple faces, Sync Labs' default behavior selects the most prominent face in the frame. To target a specific person, use the [Speaker Selection](/developer-guides/speaker-selection) feature. Speaker selection lets you identify the correct face using automatic detection (`auto_detect`) or by providing a bounding box or frame number reference. In the API, set the `active_speaker_detection` option in your generation request. In Sync Labs Studio, use the speaker selection tool in the video player controls. Note that speaker selection is available for lipsync models only -- [react-1](/models/react) does not support this feature and requires a single visible speaker. For videos with multiple speakers taking turns, use the [Segments API](/developer-guides/segments) to define time ranges and assign speaker selection per segment.

#### Output has a watermark

Watermarks appear on videos generated with free or Hobbyist accounts. To remove watermarks, upgrade to the Creator plan or higher -- watermark removal is included on all Creator+ plans. Upgrading also removes the watermark from videos you already generated on a free or Hobbyist plan. The output updates to the unwatermarked version automatically after you upgrade, with no need to regenerate. In Sync Labs Studio, refresh the page to see the change. Videos generated before October 8, 2025 may still need to be regenerated to remove the watermark. See the [Billing](/docs/product/billing) page for plan details, pricing, and upgrade instructions.

## Quality Optimization Reference

### Input Video Checklist

* Front-facing or near-frontal camera angle
* Even lighting on face, no harsh shadows
* No obstructions covering mouth (hands, microphones, hair, sunglasses) — or use sync-3 which handles obstructions automatically
* Stable camera (tripod recommended)
* Single face in frame (or use speaker selection for multi-face)
* Minimum 480p resolution, recommended 1080p, max 4K
* MP4 with H.264 codec preferred
* For color-sensitive compositing: SDR BT.709, explicit color metadata, and H.264 `yuv444p` (not `yuv420p` or `yuv422p`)

### Input Audio Checklist

* Clear speech without background music or noise
* Single speaker per audio track
* WAV or MP3 format recommended
* Duration close to video duration (use sync\_mode for mismatches)
* Any language supported

### Model Selection Guide

| Priority              | Model           | Cost/sec      |
| --------------------- | --------------- | ------------- |
| Most powerful         | `sync-3`        | $0.107-$0.133 |
| Premium detail        | `lipsync-2-pro` | $0.067-$0.083 |
| Best balance          | `lipsync-2`     | $0.04-$0.05   |
| Emotional expressions | `react-1`       | Max 15s input |

### Common Issues Quick Reference

| Issue                              | Primary Fix                                                      |
| ---------------------------------- | ---------------------------------------------------------------- |
| Lip-audio mismatch                 | Check sync\_mode, clean audio, try cut\_off mode                 |
| Long video quality drop            | Split into under 2min segments, use lipsync-2-pro                |
| Face distortion/artifacts          | Better lighting, front-facing angle, try sync-3 or lipsync-2-pro |
| Wrong face synced                  | Use active\_speaker\_detection or bounding box                   |
| Watermark on output                | Upgrade to Creator plan or higher                                |
| Poor quality on AI-generated faces | Add "person is speaking naturally" to generation prompt          |