> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server. # Generation Status Update POST Receive a notification when a generation changes status Reference: https://sync.so/docs/api-reference/api/webhooks-payload-reference/webhooks/generation-status-update ## Request ### Headers - `X-Signature-Primary` (string, required) — An HMAC signature of the payload ### Payload - `status` (enum, required) — The status of the generation - Allowed values: `COMPLETED`, `FAILED` - `createdAt` (datetime, required) — The date and time the generation was created - `id` (string, required) — A unique identifier for the generation. - `input` (list of Input, required) — The input items for the generation - `model` (enum, required) — name of the model to use for generation. - Allowed values: `sync-3`, `lipsync-2`, `lipsync-1.9.0-beta`, `lipsync-2-pro`, `react-1` - `error` (string, optional) — error message if the generation failed - `error_code` (string, optional) — error code if the generation failed - `options` (GenerationOptions, optional) — options used for the generation - `outputDuration` (double, optional) — generated output media duration in seconds - `outputUrl` (string, optional) — url of the generated output media - `segmentOutputUrl` (string, optional) — url of the segment output media - `webhookUrl` (string, optional) — url of the webhook server ## Types ### Input An input item for a generation. ### GenerationOptions - `sync_mode` (enum, optional, default: bounce) — Defines how to handle duration mismatches between video and audio inputs. Ignored for image inputs (images have no intrinsic duration). See the [Sync Mode](/developer-guides/sync-mode) guide for the full behavior matrix. - Allowed values: `bounce`, `loop`, `cut_off`, `silence`, `remap` - `model_mode` (enum, optional, default: face) — edit region for the model. only works with react-1. defaults to face, which affects lipsync + emotions in the face region. Available options are lips/face/head. When head is selected, model generates natural talking head movements along with emotions + lipsync. - Allowed values: `lips`, `face`, `head` - `prompt` (string, optional, default: neutral) — Prompt for the generation. React-1 accepts emotion prompts; the appearance model accepts a free-form appearance edit instruction. - `prompt_image_uris` (list of string, optional) — Reference image URLs for appearance editing generations. - `i2v_prompt` (string, optional) — Prompt for image-to-video generation. - `temperature` (double, optional, default: 0.5) — option to control how expressive lipsync can be. 0 -> least expressive, 1 -> most expressive. default:0.5 - `active_speaker_detection` (ActiveSpeaker, optional) — Active speaker detection configuration. When enabled, automatically detects and applies lipsync only to the active speaker in videos with multiple people. Not supported for image inputs. - `face_boxes_url` (string, optional) — URL for precomputed face bounding boxes. - `refinement_enabled` (boolean, optional) — Whether to enable the refinement pass for the generation. - `blending_mode` (enum, optional) — Controls how generated frames blend into the source media. - Allowed values: `default`, `advanced`, `disabled` - `occlusion_detection_enabled` (boolean, optional, default: false) — Whether to detect occlusion during generation, slows down generation speed. - `output_format` (enum, optional, default: mp4, deprecated) — Deprecated output container setting; defaults to mp4. - Allowed values: `mp4`, `mov` - `fps` (double, optional, deprecated) — Deprecated output frame-rate setting. - `output_resolution` (list of double, optional, deprecated) — Deprecated output resolution setting, as exactly [width, height]. Each value must be finite and between 180 and 4096 inclusive. Invalid values are discarded and the option is treated as omitted. ### Video Video input for generation. Provide either `url` or `assetId` (one is required). - `type` ("video", required) - `refId` (string, optional) — Optional reference identifier for segment definitions. Use this when a segment needs to refer back to a specific visual input item. - `url` (string, optional) — URL of the video to be used for generation. Either `url` or `assetId` must be provided. - `assetId` (string, optional) — ID of a video asset from your media library. Either `url` or `assetId` must be provided. - `segments_secs` (list of list of double, optional, deprecated) — [DEPRECATED] Use the top-level [segments](/api-reference/api/generate-api/create#request.body.segments) array instead for multi-segment support. - `segments_frames` (list of list of integer, optional, deprecated) — [DEPRECATED] Use the top-level [segments](/api-reference/api/generate-api/create#request.body.segments) array instead for multi-segment support. frames 100 and 200 of the video ### Image Image input for sync-3 model. Use instead of video when generating from a static image. Provide either `url` or `assetId` (one is required). - `type` ("image", required) - `refId` (string, optional) — Optional reference identifier for segment definitions. Use this when a segment needs to refer back to a specific visual input item. - `url` (string, optional) — URL of the image to be used for generation. Either `url` or `assetId` must be provided. - `assetId` (string, optional) — ID of an image asset from your media library. Either `url` or `assetId` must be provided. ### Audio Recorded/Captured audio input - `type` ("audio", required) - `url` (string, optional) — URL of the audio to be used for generation. Either `url` or `assetId` must be provided. - `assetId` (string, optional) — ID of an audio asset from your media library. Either `url` or `assetId` must be provided. - `refId` (string, optional) — Reference identifier for this audio input, used to link audio inputs to specific segments when using [segments](/api-reference/api/generate-api/create#request.body.segments). Required when using segments array. ### TTS Text to speech input - `type` ("text", required) - `provider` (TTSProviderConfig, required) — Integration provider configuration - `refId` (string, optional) — Reference identifier for this audio input, used to link audio inputs to specific segments when using [segments](/api-reference/api/generate-api/create#request.body.segments). Required when using segments array. ### ActiveSpeaker Active speaker detection configuration - `auto_detect` (boolean, optional, default: false) — Whether to automatically detect and apply generation to the active speaker - `v3` (boolean, optional) — Whether to use ASD v3 - `frame_number` (integer, optional) — Frame index that corresponds to the provided coordinates for manual speaker selection - `coordinates` (list of integer, optional) — Pixel coordinates [x, y] in the source video frame identified by frame_number. They are forwarded as-is to active speaker selection; they are not normalized ratios. - `bounding_boxes` (list of list of integer, optional) — Per-frame array of bounding boxes [x1, y1, x2, y2] for the detected face, or null if no box for that frame. Use instead of frame_number + coordinates when you already have detection data. - `bounding_boxes_url` (string, optional) — URL to a JSON file containing bounding boxes. Use instead of inline bounding_boxes to avoid large payloads. The JSON must have a "bounding_boxes" array with one entry per frame. - `face_image` (string, optional) — Base64-encoded reference face image (128x128 WebP) for selected-speaker detection. ### TTSProviderConfig ### ElevenLabs - `name` ("elevenlabs", required) - `voiceId` (string, required) — sync voice id (copied from cloned voices in the Studio) or ElevenLabs voice ID. Required. - `script` (string, required) — script to be used for generation - `stability` (double, optional, default: 0.5) — determines how stable the voice is and the randomness between each generation. lower values introduce broader emotional range for the voice. higher values can result in a monotonous voice with limited emotion. - `similarityBoost` (double, optional, default: 0.75) — determines how closely the ai should adhere to the original voice when attempting to replicate it.