Transcribe a public audio or video file URL directly to SRT and VTT

Last updated:

Media URL Transcriber accepts a direct public audio or video file URL and returns subtitle strings in srt and vtt, avoiding a separate download-and-upload step in your client. Use this URL-first batch workflow for public media files, or choose the Whisper-compatible endpoint when your application already submits multipart transcription requests.

Verified on October 4, 2026 against the Actor's input parser, downloader, decoder, subtitle renderer, billing loop, README and pricing configuration, plus the linked Apify and faster-whisper documentation.

A public media file URL is required, not a streaming-platform page or protected content. The input screening uses HTTP(S) and a known-platform host denylist; it is not a general allowlist of safe media hosts.

json
{
  "urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
  "language": "en",
  "model": "small",
  "outputFormats": ["text", "segments", "srt", "vtt"],
  "maxMinutesPerFile": 60
}

The example uses a public Wikimedia audio fixture already checked in the Actor's verification record. For your own media, use storage, CDN or podcast enclosure links you are allowed to process. Files need a decodable audio track: an MP4 without audio cannot produce subtitles.

InputCurrent behavior and limit
urls1–500 URLs per run, processed sequentially
File sizeAt most 1 GiB (1,073,741,824 downloaded bytes) per file
Media formatsDecodable MP3, MP4, M4A, WAV, WebM, OGG, FLAC, MKV and other PyAV-supported media; extension alone does not establish support
modelsmall by default; base is the faster, less accurate option
languageEmpty, auto or detect means detection per file; otherwise provide a supported Whisper language code such as en or ar
outputFormatsAny of text, segments, srt, vtt; all four by default
maxMinutesPerFileInteger minutes, 1–180; default 60, from the start of each file

The models cover Whisper's 99-language set; short, noisy and mixed-language recordings still need validation. Setting the known language can help, but input syntax validation does not guarantee that an arbitrary two- or three-letter code is supported by the model.

Start a run, then save subtitle text

The Actor is batch-only. Start it through the Apify run API and retrieve its dataset after completion; it does not expose an OpenAI-compatible Standby transcription endpoint.

This Python example follows the Actor's README and uses apify-client:

python
import os
from pathlib import Path
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("dropin-apis/media-url-transcriber").call(run_input={
    "urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
    "model": "small",
    "language": "en",
    "outputFormats": ["srt", "vtt"],
    "maxMinutesPerFile": 1,
})
for index, row in enumerate(client.dataset(run["defaultDatasetId"]).iterate_items(), 1):
    if row.get("error"):
        raise RuntimeError(row["error"])
    if row["truncated"]:
        raise RuntimeError("Duration cap reached; review before publishing subtitles")
    Path(f"subtitles-{index}.srt").write_text(row["srt"], encoding="utf-8")
    Path(f"subtitles-{index}.vtt").write_text(row["vtt"], encoding="utf-8")

SRT and VTT are UTF-8 strings in each dataset row, not file attachments. The renderer numbers SRT cues and uses comma milliseconds; VTT begins with WEBVTT and uses dot milliseconds. Both derive from the same segment timings.

Result fieldMeaning
url, model, languageInput source, selected model and returned language
durationSecSource duration when known; falls back to decoded duration if unavailable
transcribedSec, truncatedDuration actually decoded for transcription, and whether the cap cut the source
billedMinutesStarted minutes billed for the decoded/transcribed portion, minimum one on success
text, segments, srt, vttPresent only for the requested formats; segment objects have start, end, text
errornull on success; a handled file failure produces an error row with billedMinutes: 0

Timestamp quality and language limits

This implementation produces segment timestamps, not word timestamps, speaker labels or translations. It uses open-source faster-whisper with voice-activity filtering and disables conditioning on previous text to reduce repetition. Those choices do not guarantee correct words or perfectly synchronized captions.

Before publishing, compare several cues against playback, including the opening, the duration cap and a noisy passage. An empty set of cues can be a valid result for silence; it does not prove a spoken file was understood. base and small can disagree, and timings need review for accessibility or editing workflows.

For an existing transcription client, see keeping SRT/VTT during a Whisper migration. That is a different interface; do not send response_format to this Actor when its input field is outputFormats.

Download and safety boundaries

The initial URL rejects non-HTTP(S) schemes and known video/music-platform domains, including YouTube, TikTok, Instagram, Vimeo, Twitch, SoundCloud and Spotify. HTML responses, empty downloads, oversized files and undecodable media fail instead of being treated as transcripts.

The current downloader follows Python's default urlopen redirects and does not add public-IP DNS pinning or destination validation for each redirect. It therefore does not provide a complete private-network/SSRF barrier: use trusted direct media links, and enforce destination rules outside the Actor before accepting arbitrary third-party URLs. This is a code-observed boundary, not a claimed platform security guarantee. Python urllib.request behavior

Download operations use a 60-second socket-operation timeout with up to four attempts for retryable failures; that is not a 60-second total-file deadline. The default whole-run timeout is 10,800 seconds. Long batches may need smaller runs, particularly with small on CPU.

The downloaded temporary file is deleted after processing; transcripts remain in your Apify dataset. Treat transcript text as untrusted content in downstream agents. The Document to Markdown workflow is the adjacent route for public written documents rather than speech.

Price and truncation

ModelEventPrice per started minuteTen-minute media charge
smallaudio-minute$0.014$0.14
baseaudio-minute-base$0.006$0.06

Apify's run-start charge is additional. A 30-second successful file incurs one minute event, not half an event. Failed files have no audio-minute event; the run-start charge can still apply. After decoding, the Actor checks whether the budget can cover the next file before transcription and stops if it cannot.

The duration cap limits decoded/transcribed audio, not the download: the file is downloaded before truncation. Always inspect truncated and transcribedSec; a successful run can contain only the beginning of a longer recording, or fewer results because the spending cap ended the batch.

Run Media URL Transcriber on Apify with a short known file and inspect the subtitle strings first.

FAQ

Can I transcribe a YouTube or TikTok page URL?

No. Streaming-platform pages are unsupported. Supply a direct public audio or video file URL you are allowed to process.

Does this produce word timestamps or translate into English?

No. It returns segment timings and a transcript in the spoken language. There is no word-timestamp, speaker-diarization or translation option in this Actor.

Why are my subtitles shorter than the recording?

maxMinutesPerFile defaults to 60 and can be set from 1 to 180. Inspect truncated and transcribedSec; the Actor transcribes only the beginning up to that cap.

Are SRT and VTT downloadable files?

They are strings in dataset rows. Save the srt and vtt fields as UTF-8 files, as shown above, after checking for an error or truncation.