Transcribe an audio or video URL to text, SRT and VTT

Last updated:

Give this Actor a direct URL to an audio or video file and it returns the transcript, timed segments, and ready-to-save SRT and VTT subtitles. It runs open-source Whisper (faster-whisper) in 99 languages with automatic language detection, and bills per started minute of audio — no subscription.

What it accepts

mp3, mp4, m4a, wav, webm, ogg, flac, mkv and more — any direct link to a media file you're allowed to process (your own storage, S3, a CDN, a podcast enclosure URL). It does not download from YouTube, TikTok, Instagram or other platform pages — pass a direct file link.

Example input

json
{
  "urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
  "language": "",
  "model": "small",
  "outputFormats": ["text", "segments", "srt", "vtt"],
  "maxMinutesPerFile": 60
}

Two models: small (default, more accurate, $0.014/min) and base (2–3× faster, less accurate, $0.006/min). Up to 500 URLs and 1 GB per file; maxMinutesPerFile (1–180) caps cost by cutting long files at that point and billing only the transcribed minutes.

Example output (real run on Apify)

json
{
  "url": "https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg",
  "language": "en",
  "durationSec": 24.11,
  "billedMinutes": 1,
  "text": "I'm going to step off the land now. That's one small step for man. One giant leap for mankind.",
  "segments": [
    { "start": 3.38, "end": 15.18, "text": "I'm going to step off the land now." },
    { "start": 15.18, "end": 20.81, "text": "That's one small step for man." }
  ],
  "srt": "1\n00:00:03,380 --> 00:00:15,180\nI'm going to step off the land now.\n\n...",
  "vtt": "WEBVTT\n\n00:00:03.380 --> 00:00:15.180\nI'm going to step off the land now.\n\n...",
  "error": null
}

A file that can't be processed (404, a web page instead of a media file, no audio track) returns a row with error set and billedMinutes: 0 — the run continues with the next file.

How it compares

Audiosmall modelbase model
30-second voice note$0.014$0.006
10-minute speech$0.14$0.06
1-hour podcast$0.84$0.36

Measured on Apify's default 2 GB memory, a 10.3-minute speech took 7 minutes with small; at 4 GB it took under 4 minutes, and base took 1.5 minutes. Accuracy is that of the open Whisper small model: good on clear speech, weaker than large commercial models on noisy audio, heavy accents and rare languages.

Run it from code

python
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("dropin-apis/media-url-transcriber").call(run_input={
    "urls": ["https://example.com/podcast/episode-12.mp3"],
    "outputFormats": ["text", "srt"],
})

Agents connected to the Apify MCP server can find it with search-actors ("transcribe audio") and run it with call-actor.

Pricing

Pay per started minute of transcribed audio, plus Apify's standard run-start charge. $0.014/minute (small, default) or $0.006/minute (base). No subscription; failed files are not charged, and a maximum-charge budget on the run is checked before each file, stopping cleanly rather than over-billing.

Open the Whisper Audio & Video URL Transcriber on Apify

FAQ

How do I transcribe an MP4 video to text?

Pass the direct URL of the .mp4 file in urls. The Actor extracts the audio track automatically and returns the transcript, segments, SRT and VTT.

Can I make SRT or VTT subtitles from an audio file?

Yes — include srt and/or vtt in outputFormats (both are included by default), and each result row has ready-to-save subtitle text.

No. The Actor does not download from video or music platforms. Use a direct link to a media file you have the right to process.

Which languages are supported?

All 99 Whisper languages, including English, Arabic, Spanish, French, German, Hindi, Chinese and Japanese, detected per file unless you set language.

What does it cost?

$0.014 per started minute with the default small model ($0.84/hour), or $0.006/minute with base ($0.36/hour). Failed files are not charged.

Is my audio stored anywhere?

No. Each file is downloaded to a temporary file inside the run, transcribed, and deleted. Only the transcript is saved to your dataset.

I need an OpenAI whisper-1 compatible endpoint instead

Use the sister Actor, the whisper-1 alternative — the same engine behind OpenAI's /v1/audio/transcriptions request and response format, for apps already built on the OpenAI SDK.