> For the documentation index, fetch https://sync.so/docs/llms.txt. Append .md to a page URL for Markdown. Documentation-search MCP: https://sync.so/docs/_mcp/server.

# Generation Status Update

POST 

Receive a notification when a generation changes status

Reference: https://sync.so/docs/api-reference/api/webhooks-payload-reference/webhooks/generation-status-update

## Request

### Headers

- `X-Signature-Primary` (string, required) — An HMAC signature of the payload

### Payload

- `status` (enum, required) — The status of the generation
  - Allowed values: `COMPLETED`, `FAILED`
- `createdAt` (datetime, required) — The date and time the generation was created
- `id` (string, required) — A unique identifier for the generation.
- `input` (list of Input, required) — The input items for the generation
- `model` (enum, required) — name of the model to use for generation.
  - Allowed values: `sync-3`, `lipsync-2`, `lipsync-1.9.0-beta`, `lipsync-2-pro`, `react-1`
- `error` (string, optional) — error message if the generation failed
- `error_code` (string, optional) — error code if the generation failed
- `options` (GenerationOptions, optional) — options used for the generation
- `outputDuration` (double, optional) — generated output media duration in seconds
- `outputUrl` (string, optional) — url of the generated output media
- `segmentOutputUrl` (string, optional) — url of the segment output media
- `webhookUrl` (string, optional) — url of the webhook server

## Types

### Input

An input item for a generation.

### GenerationOptions

- `sync_mode` (enum, optional, default: bounce) — Defines how to handle duration mismatches between video and audio inputs. Ignored for image inputs (images have no intrinsic duration). See the [Sync Mode](/developer-guides/sync-mode) guide for the full behavior matrix.
  - Allowed values: `bounce`, `loop`, `cut_off`, `silence`, `remap`
- `model_mode` (enum, optional, default: face) — edit region for the model. only works with react-1. defaults to face, which affects lipsync + emotions in the face region. Available options are lips/face/head. When head is selected, model generates natural talking head movements along with emotions + lipsync.
  - Allowed values: `lips`, `face`, `head`
- `prompt` (string, optional, default: neutral) — Prompt for the generation. React-1 accepts emotion prompts; the appearance model accepts a free-form appearance edit instruction.
- `prompt_image_uris` (list of string, optional) — Reference image URLs for appearance editing generations.
- `i2v_prompt` (string, optional) — Prompt for image-to-video generation.
- `temperature` (double, optional, default: 0.5) — option to control how expressive lipsync can be. 0 -> least expressive, 1 -> most expressive. default:0.5
- `active_speaker_detection` (ActiveSpeaker, optional) — Active speaker detection configuration. When enabled, automatically detects and applies lipsync only to the active speaker in videos with multiple people. Not supported for image inputs.
- `face_boxes_url` (string, optional) — URL for precomputed face bounding boxes.
- `refinement_enabled` (boolean, optional) — Whether to enable the refinement pass for the generation.
- `blending_mode` (enum, optional) — Controls how generated frames blend into the source media.
  - Allowed values: `default`, `advanced`, `disabled`
- `occlusion_detection_enabled` (boolean, optional, default: false) — Whether to detect occlusion during generation, slows down generation speed.
- `output_format` (enum, optional, default: mp4, deprecated) — Deprecated output container setting; defaults to mp4.
  - Allowed values: `mp4`, `mov`
- `fps` (double, optional, deprecated) — Deprecated output frame-rate setting.
- `output_resolution` (list of double, optional, deprecated) — Deprecated output resolution setting, as exactly [width, height]. Each value must be finite and between 180 and 4096 inclusive. Invalid values are discarded and the option is treated as omitted.

### Video

Video input for generation. Provide either `url` or `assetId` (one is required).

- `type` ("video", required)
- `refId` (string, optional) — Optional reference identifier for segment definitions. Use this when a segment needs to refer back to a specific visual input item.
- `url` (string, optional) — URL of the video to be used for generation. Either `url` or `assetId` must be provided.
- `assetId` (string, optional) — ID of a video asset from your media library. Either `url` or `assetId` must be provided.
- `segments_secs` (list of list of double, optional, deprecated) — [DEPRECATED] Use the top-level [segments](/api-reference/api/generate-api/create#request.body.segments) array instead for multi-segment support.
- `segments_frames` (list of list of integer, optional, deprecated) — [DEPRECATED] Use the top-level [segments](/api-reference/api/generate-api/create#request.body.segments) array instead for multi-segment support. frames 100 and 200 of the video

### Image

Image input for sync-3 model. Use instead of video when generating from a static image. Provide either `url` or `assetId` (one is required).

- `type` ("image", required)
- `refId` (string, optional) — Optional reference identifier for segment definitions. Use this when a segment needs to refer back to a specific visual input item.
- `url` (string, optional) — URL of the image to be used for generation. Either `url` or `assetId` must be provided.
- `assetId` (string, optional) — ID of an image asset from your media library. Either `url` or `assetId` must be provided.

### Audio

Recorded/Captured audio input

- `type` ("audio", required)
- `url` (string, optional) — URL of the audio to be used for generation. Either `url` or `assetId` must be provided.
- `assetId` (string, optional) — ID of an audio asset from your media library. Either `url` or `assetId` must be provided.
- `refId` (string, optional) — Reference identifier for this audio input, used to link audio inputs to specific segments when using [segments](/api-reference/api/generate-api/create#request.body.segments). Required when using segments array.

### TTS

Text to speech input

- `type` ("text", required)
- `provider` (TTSProviderConfig, required) — Integration provider configuration
- `refId` (string, optional) — Reference identifier for this audio input, used to link audio inputs to specific segments when using [segments](/api-reference/api/generate-api/create#request.body.segments). Required when using segments array.

### ActiveSpeaker

Active speaker detection configuration

- `auto_detect` (boolean, optional, default: false) — Whether to automatically detect and apply generation to the active speaker
- `v3` (boolean, optional) — Whether to use ASD v3
- `frame_number` (integer, optional) — Frame index that corresponds to the provided coordinates for manual speaker selection
- `coordinates` (list of integer, optional) — Pixel coordinates [x, y] in the source video frame identified by frame_number. They are forwarded as-is to active speaker selection; they are not normalized ratios.
- `bounding_boxes` (list of list of integer, optional) — Per-frame array of bounding boxes [x1, y1, x2, y2] for the detected face, or null if no box for that frame. Use instead of frame_number + coordinates when you already have detection data.
- `bounding_boxes_url` (string, optional) — URL to a JSON file containing bounding boxes. Use instead of inline bounding_boxes to avoid large payloads. The JSON must have a "bounding_boxes" array with one entry per frame.
- `face_image` (string, optional) — Base64-encoded reference face image (128x128 WebP) for selected-speaker detection.

### TTSProviderConfig

### ElevenLabs

- `name` ("elevenlabs", required)
- `voiceId` (string, required) — sync voice id (copied from cloned voices in the Studio) or ElevenLabs voice ID. Required.
- `script` (string, required) — script to be used for generation
- `stability` (double, optional, default: 0.5) — determines how stable the voice is and the randomness between each generation. lower values introduce broader emotional range for the voice. higher values can result in a monotonous voice with limited emotion.
- `similarityBoost` (double, optional, default: 0.75) — determines how closely the ai should adhere to the original voice when attempting to replicate it.