Zur Navigation springen

Transcription and Dialogue Editing

The dialogue workflow has three steps: transcribe the source, preview changes to its words, then create a lip sync generation using that preview. Transcription and preview jobs are asynchronous. Keep their IDs so your application can resume polling after a disconnect.

Prepare the source

Transcription and dialogue edits require media hosted in Sync storage. For local or externally hosted files, follow Asset Uploads: request an upload URL, upload the bytes, and register the asset. Keep the hosted source video URL and asset ID.

Transcription accepts exactly one of sourceVideoUrl and sourceAudioUrl. Supplying extracted audio avoids server-side extraction. A dialogue edit requires the source video URL and word timings on that video’s timeline. Use the video’s own transcription for this flow.

Transcribe the video

curl https://api.sync.so/v2/transcriptions \
-H "x-api-key: $SYNC_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"sourceVideoUrl": "https://assets.sync.so/docs/example-video.mp4"}'

Submission returns 201 with a job. Poll GET /v2/transcriptions/{id} using the same API key until status is COMPLETED or FAILED. Use backoff between requests. On completion, save the entire transcript, including segment and word IDs, timing, and speaker information. speakerCount and optional difficulty help describe the source; difficulty is informational.

Word startMs and endMs values are absolute milliseconds from the beginning of the source. Optional request startMs defaults to zero; endMs must be greater than startMs. An explicit transcription window can be at most 60 minutes. maxSourceSeconds, when supplied, is positive and cannot exceed 3,600. Identical source-and-window requests reuse an existing job. Transcription does not charge generation credits.

Preview an edit

Send the returned transcript unchanged and describe the edits separately. An operation with kind: "change" replaces one word by its wordId. An operation with kind: "remove" lists contiguous wordIds. Supply 1–100 operations.

The following Python example uses the actual transcript returned by a completed transcription and replaces its first word. Set TRANSCRIPTION_ID to your completed transcription job ID. Review the selected word before applying this example to your own dialogue.

import os
import requests
headers = {"x-api-key": os.environ["SYNC_API_KEY"]}
transcription_id = os.environ["TRANSCRIPTION_ID"]
response = requests.get(
f"https://api.sync.so/v2/transcriptions/{transcription_id}",
headers=headers,
timeout=30,
)
response.raise_for_status()
transcription = response.json()
if transcription["status"] != "COMPLETED":
raise RuntimeError("Wait for the transcription to complete first")
transcript = transcription["transcript"]
word = next(word for segment in transcript["segments"] for word in segment["words"])
response = requests.post(
"https://api.sync.so/v2/dialogue-edits",
headers=headers,
json={
"sourceVideoUrl": transcription["sourceVideoUrl"],
"transcript": transcript,
"edits": [{"kind": "change", "wordId": word["id"], "replacement": "Hello"}],
},
timeout=30,
)
response.raise_for_status()
print(response.json()["id"])

Submission returns 201. Poll GET /v2/dialogue-edits/{id} until COMPLETED, COMPLETED_PARTIAL, or FAILED. Both completed statuses are terminal and can carry playable preview audio. Listen to previewAudioUrl before proceeding. previewDurationMs, resultTranscript, and resultSlots describe the edited timing. previewAssetId can contain the registered preview asset ID when available. It can be null or absent for an unregistered preview.

To revise a preview, submit another job with the adjusted edits. You can pass the previous job’s voiceId to reuse its cloned voice when it was used by your organization under the same shared or connected ElevenLabs key pool. rerunOfJobId records a previous COMPLETED or COMPLETED_PARTIAL job in your organization. It is informational lineage, not a retry or idempotency key. Hints that do not meet these conditions are ignored.

API-key dialogue edits must cover the whole video: omit startMs and endMs. The whole source must fit the 10-minute dialogue limit. The organization must have both segment lipsync and section expansion enabled. Session-authenticated clients can select a window, with a maximum of 10 minutes. Billing, planning, duration and rate limits still apply.

Lip sync the approved preview

Submit dialogueEdit: { "id": "..." } with exactly one video input. It must be the original video used to create the edit. Use its URL or registered asset ID. Do not add an audio input, image input, segments, or dubParams.

curl https://api.sync.so/v2/generate \
-H "x-api-key: $SYNC_API_KEY" \
-H 'Content-Type: application/json' \
-H 'Idempotency-Key: approved-dialogue-take-001' \
-d '{
"model": "lipsync-2",
"input": [{"type": "video", "url": "https://assets.sync.so/docs/example-video.mp4"}],
"dialogueEdit": {"id": "REPLACE_WITH_COMPLETED_DIALOGUE_JOB_ID"}
}'

The completed preview must report both segmentLipsyncEnabled and sectionExpansionEnabled as true. Continue with ordinary generation polling and webhooks. Generation billing and limits apply.

If sending multipart form data, encode dialogueEdit and input as JSON strings. Uploaded video, audio and image file parts are rejected for this workflow. Uploading the video again is not a substitute for referencing the original source.

Handle refusals and failed jobs

StageFailureNext step
Transcription admission422 for a non-Sync-hosted source or SOURCE_TOO_LONGUpload to Sync storage or shorten the source/window.
Transcription admission429 for caller rate or organization daily limitBack off and respect the applicable limit.
Transcription jobNO_AUDIO, NO_AUDIO_TRACK, SOURCE_TOO_LONGFix the source; repeating the same request does not repair it.
Dialogue admission400 PLAN_FAILEDCorrect word references and edit operations.
Dialogue admission422 dialogue_edit_unsupportedFor API keys, omit the partial window.
Dialogue admission or generation422 dialogue_edit_retime_requiredThe required rollout capabilities are unavailable; do not submit the preview as ordinary audio to bypass this constraint.
Dialogue admission503 generation_infra_service_unavailable, generation_admission_paused, or generation_admission_unavailableAdmission is temporarily unavailable; retry with backoff and honor Retry-After if present.
Generation422 dialogue_edit_source_mismatchReference the original video by URL or asset ID.
Generation422 dialogue_edit_audio_conflictRemove the audio input.
Generation422 dialogue_edit_removal_too_largeReduce the removal or revise the edit. When returned, dialogueEditSection identifies the section to correct.

Billing holds can return 402; dialogue output duration is also subject to plan limits. Other transcription job failures include EXTRACTION_FAILED, PROVIDER_FAILED, TIMEOUT, and UNKNOWN. Dialogue job errors include extraction, voice cloning, short voice samples, synthesis, edit/source length, planning, timeout, and unknown failures. Inspect the job’s error.code and error.message before creating another attempt. Polling a failed job does not restart it.