Edit dialogue and preview the result as audio
Applies word-level edits to a transcript from POST /v2/transcriptions and synthesizes preview audio for the same window: only the edited spans are re-spoken in a voice cloned from the source, everything else stays the original recording. Poll GET /v2/dialogue-edits/{id} until status is COMPLETED, COMPLETED_PARTIAL or FAILED; a completed or partially completed job carries previewAudioUrl, the retimed resultTranscript and the edited regions as resultSlots. Edits that cannot be applied to the transcript are refused with a 400 before a job exists. An API-key request is admitted only for the whole video (leave startMs and endMs unset) and only when the organization’s rollout enables segment lipsync and section expansion: a partial window is refused with errorCode dialogue_edit_unsupported (422), a rollout without both flags with dialogue_edit_retime_required (422), and a rollout that cannot be evaluated with generation_infra_service_unavailable (503; retry with backoff). Studio sessions may select a window. Normal billing, planning, duration and rate limits still apply.
Autenticación
Solicitud
The transcript the edits refer to, as returned by GET /v2/transcriptions/{id}. Edits name words by their id in this transcript. Word timings are in milliseconds from the start of the source video.
The edits to apply: change replaces one word (wordId, replacement), remove cuts a contiguous run of words (wordIds). Between 1 and 100 edits per job.
URL of the source video the transcript was made from, hosted in sync. storage (upload it via POST /v2/assets/upload first).
Start of the edited window, in milliseconds from the start of the source. Use the same window the transcript was made for. Defaults to 0. An API-key request must leave startMs and endMs unset; the edit then covers the whole video.
End of the edited window, in milliseconds from the start of the source. Omit to run to the end. The window cannot exceed 10 minutes. An API-key request must leave startMs and endMs unset to edit the whole video.
Opaque voice id returned as voiceId by an earlier dialogue-edit job of the same source. Reusing it skips cloning the voice again. Omit on the first job for a source. Ids your organization has not used before are ignored.
Id of the completed dialogue-edit job this request revises, when regenerating a preview after adjusting the edits. Informational lineage only; unknown ids are ignored.
Respuesta
Job id. Poll GET /v2/dialogue-edits/{id} with it.
Job status. COMPLETED, COMPLETED_PARTIAL and FAILED are terminal. COMPLETED_PARTIAL means the preview is playable but shipped without something that was requested; treat it as completed.
Whether the preview was cut for per-region lipsync (true) or for re-lipsyncing the whole window (false). Decided when the job is created and fixed for its lifetime. For jobs created with an API key, the rollout is evaluated against the organization's members.
Whether the preview was made under the section-expansion policy. Decided when the job is created and fixed for its lifetime. For jobs created with an API key, the rollout is evaluated against the organization's members.
Hosted WAV of the edited audio for the whole window: the original recording with only the edited spans re-synthesized. Null until the job completes.
Asset id of the registered preview after completion, when available. Null or absent before completion or for an unregistered preview. To lipsync the edit, pass the job id as dialogueEdit: { id } on POST /v2/generate.
The edited regions in order, each located on the preview timeline (outputStartMs, outputDurationMs) and in the source video (sourceStartMs, sourceDurationMs). Everything between two slots is untouched source. Null until the job completes.
Opaque id of the voice cloned from the source and used for synthesis. Pass it as voiceId on the next dialogue-edit job of the same source to skip cloning. Null until the voice exists.

