Transcribe a video for dialogue editing
Starts a word-level transcription of a window of a video, the first step of a dialogue edit. Provide exactly one of sourceVideoUrl or sourceAudioUrl, both hosted in sync. storage (upload via POST /v2/assets/upload). Passing pre-extracted audio skips the server-side download and transcode. Poll GET /v2/transcriptions/{id} until status is COMPLETED, then pass the transcript and your edits to POST /v2/dialogue-edits. An identical source and window returns the existing job instead of starting a new one. Transcription is not billed.
Authentifizierung
Anfrage
URL of the source video, hosted in sync. storage (upload it via POST /v2/assets/upload first). Provide exactly one of sourceVideoUrl or sourceAudioUrl.
URL of audio already extracted from the source, hosted in sync. storage. Skips the server-side audio extraction. Provide exactly one of sourceVideoUrl or sourceAudioUrl.
Refuse the source if it is longer than this many seconds. Useful when a later step, such as a dialogue edit (10 minute window), cannot consume a longer transcript. Cannot exceed the endpoint ceiling of 3600 seconds.
Antwort
Job id. Poll GET /v2/transcriptions/{id} with it.
Word-level transcript. Word timings are in milliseconds from the start of the source video, not the window. Null until status is COMPLETED.
Number of distinct speakers detected. Dialogue edits currently support a single speaker; a count above 1 completes the job but a dialogue edit of that source is unsupported. Null until status is COMPLETED.

