> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server. # Concurrency & Rate Limiting > Understanding sync. labs API rate limits, concurrency limits, and retry strategies for lip sync generation. ## Rate Limiting Request-rate limits count requests over a time window. API-key requests are limited per key; signed-in user requests are limited per user, and unauthenticated requests fall back to the caller's IP address. These limits are separate from generation concurrency. ### API Endpoints | Endpoint | Request-rate limit | | :----------------------- | :----------------------------------------------- | | `POST /v2/tts` | 60 requests/min | | `POST /v2/tts/jobs` | 60 requests/min, separately from synchronous TTS | | `POST /v2/voices` | 10 requests/hour | | `POST /v2/assets/upload` | 120 requests/min | Generation creation also enforces the plan's concurrency limit described below. Inspect the error body instead of treating every `429` as the same limit. ### Auth Endpoints | Endpoint | Limit | | :----------------------------------------------------------------------------- | :-------------- | | `POST /api/auth/sign-in/*`, `POST /api/auth/sign-up/*`, `POST /api/auth/otp/*` | 10 requests/min | | `GET /api/auth/session` | 60 requests/min | When rate limited, the API returns a `429` response: ```json { "statusCode": 429, "errorCode": "rate_limit_exceeded", "message": "Rate limit exceeded. Please try again later.", "error": "Too Many Requests" } ``` ## Concurrency Concurrency refers to the number of generations that can be submitted/processed concurrently. Requests to create new generations will fail with a 429 error if the concurrency limit is exceeded. To check your generations currently in PENDING/PROCESSING state, you can use the [List Generations](/api-reference/api/generate-api/list) endpoint. Concurrency limits are defined in the subscription plan. For credit-billing plans, see the [current credit plan limits](/product/billing/v3). The legacy subscription limits are: | Plan | Concurrent Requests | | ---------- | ------------------- | | Free | 1 | | Hobbyist | 1 | | Creator | 3 | | Growth | 6 | | Scale | 15 | | Enterprise | Custom | > **Note** > > Enterprise and contracted accounts may have custom concurrency that is higher than the public plan table. If your contract or support contact gave you a custom limit, use that value instead of hardcoding the public defaults. ## Handling 429 Errors When you exceed a rate limit or concurrency limit, the API returns a `429 Too Many Requests` response. This applies to both per-minute rate limits and concurrent generation limits. Changing your IP address does not bypass limits tied to your API key or user account. What to do when you hit a 429: * **Rate limit (requests per minute):** Wait briefly and retry. Honor the response's `retryAfter` and `resetTime` values when present; endpoint windows differ. * **Concurrency limit:** You already have the maximum number of generations in PENDING or PROCESSING state for your plan. Wait for an existing generation to complete before submitting a new one, or upgrade your plan for higher limits. Do not retry 429 responses immediately in a tight loop. This wastes requests and delays recovery. Use the retry strategy below instead. ## Retry Strategies Exponential backoff is the recommended approach for handling both rate limit and transient errors from the Sync Labs lip sync API. Each retry waits progressively longer, reducing pressure on the API and improving your success rate. #### Python ```python import time from sync import Sync from sync.common import Audio, Video from sync.core.api_error import ApiError sync = Sync() def create_with_retry(video_url: str, audio_url: str, max_retries: int = 5): for attempt in range(max_retries): try: return sync.generations.create( input=[Video(url=video_url), Audio(url=audio_url)], model="lipsync-2", ) except ApiError as e: if e.status_code == 429 and attempt < max_retries - 1: wait = 2 ** attempt # 1s, 2s, 4s, 8s, 16s print(f"Rate limited. Retrying in {wait}s...") time.sleep(wait) else: raise response = create_with_retry( "https://assets.sync.so/docs/example-video.mp4", "https://assets.sync.so/docs/example-audio.wav", ) print(f"Job submitted: {response.id}") ``` #### TypeScript ```typescript import { SyncClient, SyncError } from "@sync.so/sdk"; const sync = new SyncClient(); async function createWithRetry(videoUrl: string, audioUrl: string, maxRetries = 5) { for (let attempt = 0; attempt < maxRetries; attempt++) { try { return await sync.generations.create({ input: [ { type: "video", url: videoUrl }, { type: "audio", url: audioUrl }, ], model: "lipsync-2", }); } catch (err) { if (err instanceof SyncError && err.statusCode === 429 && attempt < maxRetries - 1) { const wait = 2 ** attempt * 1000; // 1s, 2s, 4s, 8s, 16s console.log(`Rate limited. Retrying in ${wait / 1000}s...`); await new Promise((r) => setTimeout(r, wait)); } else { throw err; } } } } const response = await createWithRetry( "https://assets.sync.so/docs/example-video.mp4", "https://assets.sync.so/docs/example-audio.wav", ); console.log(`Job submitted: ${response.id}`); ``` ## Optimizing Concurrency To get the most out of your plan's concurrent generation slots: * **Monitor active jobs.** Use the [List Generations](/api-reference/api/generate-api/list) endpoint to check how many jobs are currently in PENDING or PROCESSING state before submitting new ones. * **Queue on your side.** Maintain a local queue and only submit a new generation when a slot frees up. This avoids wasted 429 responses. * **Use webhooks.** Configure a [webhook](/api-reference/guides/webhooks) to get notified when a generation completes, so you can immediately submit the next job without polling. * **Batch when possible.** If you are on a Scale or Enterprise plan, the [Batch API](/api-reference/guides/batch-processing) handles queueing and concurrency for you -- submit up to 500 generations in one call. For API rate limiting best practices and general error handling, see the [Error Handling](/developer-guides/error-handling) guide. > Understanding rate limiting and concurrency