Transcription and Dialogue Editing
The dialogue workflow has three steps: transcribe the source, preview changes to its words, then create a lip sync generation using that preview. Transcription and preview jobs are asynchronous. Keep their IDs so your application can resume polling after a disconnect.
Prepare the source
Transcription and dialogue edits require media hosted in Sync storage. For local or externally hosted files, follow Asset Uploads: request an upload URL, upload the bytes, and register the asset. Keep the hosted source video URL and asset ID.
Transcription accepts exactly one of sourceVideoUrl and sourceAudioUrl. Supplying extracted audio avoids server-side extraction. A dialogue edit requires the source video URL and word timings on that video’s timeline. Use the video’s own transcription for this flow.
Transcribe the video
Submission returns 201 with a job. Poll GET /v2/transcriptions/{id} using the same API key until status is COMPLETED or FAILED. Use backoff between requests. On completion, save the entire transcript, including segment and word IDs, timing, and speaker information. speakerCount and optional difficulty help describe the source; difficulty is informational.
Word startMs and endMs values are absolute milliseconds from the beginning of the source. Optional request startMs defaults to zero; endMs must be greater than startMs. An explicit transcription window can be at most 60 minutes. maxSourceSeconds, when supplied, is positive and cannot exceed 3,600. Identical source-and-window requests reuse an existing job. Transcription does not charge generation credits.
Preview an edit
Send the returned transcript unchanged and describe the edits separately. An operation with kind: "change" replaces one word by its wordId. An operation with kind: "remove" lists contiguous wordIds. Supply 1–100 operations.
The following Python example uses the actual transcript returned by a completed transcription and replaces its first word. Set TRANSCRIPTION_ID to your completed transcription job ID. Review the selected word before applying this example to your own dialogue.
Submission returns 201. Poll GET /v2/dialogue-edits/{id} until COMPLETED, COMPLETED_PARTIAL, or FAILED. Both completed statuses are terminal and can carry playable preview audio. Listen to previewAudioUrl before proceeding. previewDurationMs, resultTranscript, and resultSlots describe the edited timing. previewAssetId can contain the registered preview asset ID when available. It can be null or absent for an unregistered preview.
To revise a preview, submit another job with the adjusted edits. You can pass the previous job’s voiceId to reuse its cloned voice when it was used by your organization under the same shared or connected ElevenLabs key pool. rerunOfJobId records a previous COMPLETED or COMPLETED_PARTIAL job in your organization. It is informational lineage, not a retry or idempotency key. Hints that do not meet these conditions are ignored.
API-key dialogue edits must cover the whole video: omit startMs and endMs. The whole source must fit the 10-minute dialogue limit. The organization must have both segment lipsync and section expansion enabled. Session-authenticated clients can select a window, with a maximum of 10 minutes. Billing, planning, duration and rate limits still apply.
Lip sync the approved preview
Submit dialogueEdit: { "id": "..." } with exactly one video input. It must be the original video used to create the edit. Use its URL or registered asset ID. Do not add an audio input, image input, segments, or dubParams.
The completed preview must report both segmentLipsyncEnabled and sectionExpansionEnabled as true. Continue with ordinary generation polling and webhooks. Generation billing and limits apply.
If sending multipart form data, encode dialogueEdit and input as JSON strings. Uploaded video, audio and image file parts are rejected for this workflow. Uploading the video again is not a substitute for referencing the original source.
Handle refusals and failed jobs
Billing holds can return 402; dialogue output duration is also subject to plan limits. Other transcription job failures include EXTRACTION_FAILED, PROVIDER_FAILED, TIMEOUT, and UNKNOWN. Dialogue job errors include extraction, voice cloning, short voice samples, synthesis, edit/source length, planning, timeout, and unknown failures. Inspect the job’s error.code and error.message before creating another attempt. Polling a failed job does not restart it.

