Does the whisper-1 retirement break SRT and VTT output?

Last updated:

⏳ 144 days left — the old API shuts down February 26, 2027.

It can. OpenAI removes whisper-1 on February 26, 2027. Its API reference lists json as the only response format for gpt-4o-transcribe and gpt-4o-mini-transcribe, and its guide documents word timestamps only for whisper-1. Test srt, vtt and timestamps on the replacement before you change the model name. If you need the exact whisper-1 request and output, point the OpenAI SDK at a whisper-1-compatible endpoint, or use a URL-to-subtitles tool when your input is a public media file.

Facts on this page were checked against OpenAI's deprecations page, speech-to-text guide and API reference, and the two Actors' READMEs, on October 3, 2026.

What OpenAI's documentation says

Developers already hit subtitle problems with whisper-1 itself, for example SRT generation and VTT timestamps that do not line up (Stack Overflow). A migration is a good time to test those outputs properly rather than assume them.

What to test before you switch

You useCheck on the replacement
response_format=srt or vttIs the format accepted? Do cue boundaries and timing match your player and caption pipeline?
verbose_json with timestamp_granularities[]=wordAre word timestamps available at all? If not, what replaces karaoke captions, search-in-audio or editing tools?
/v1/audio/translationsIs there an equivalent path to English, or do you need a separate translation step?
Language detection and the language fieldSame language codes, and the same behavior on mixed-language audio?
Error handlingDo error codes and messages still match what your code parses?

Run a handful of your real files through both before changing production. A sample that matches on clean speech can still differ on noise, music or accents.

Two ways to keep subtitles working

Your inputOptionOutputPrice
You send audio files through the OpenAI SDKwhisper-1-compatible endpoint: keep model="whisper-1", change base_url and api_keyjson, text, srt, vtt, verbose_json with word and segment timestamps; /v1/audio/translations$0.006 per minute, billed per started 15 s
You have a direct public URL to a media fileMedia URL transcriber: send the URL, get every format back in one rowtext, segments, srt, vtt (segment timestamps, not word timestamps)$0.006/min (base) or $0.014/min (small)

Both run the open Whisper models in 99 languages, not OpenAI's hosted models. Results are close to whisper-1 on clean speech but not identical:

diff
- client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
+ client = OpenAI(base_url="https://dropin-apis--whisper-compat.apify.actor/v1", api_key=os.environ["APIFY_TOKEN"])
  srt = client.audio.transcriptions.create(model="whisper-1", file=f, response_format="srt")

FAQ

Will my SRT and VTT code keep working after February 26, 2027?

Not with whisper-1 on OpenAI, because the model is removed that day. gpt-4o-transcribe and gpt-4o-mini-transcribe only return json, and are removed the same day. Test subtitle output on gpt-transcribe, or keep the whisper-1 request on a compatible endpoint.

Does gpt-transcribe support word timestamps?

OpenAI's speech-to-text guide documents timestamp_granularities[] only for whisper-1. Check the current guide and test before you rely on word timestamps with another model.

Is a whisper-1-compatible endpoint the same model as whisper-1?

No. It keeps the request and the response format. The transcription comes from open Whisper models, so the words can differ on difficult audio.