tts-1 is removed on 6 January 2027. Here's a compatible endpoint.

Last updated:

⏳ 93 days left — the old API shuts down January 6, 2027.

OpenAI is removing tts-1, tts-1-hd and both gpt-4o-mini-tts snapshots from the API on 6 January 2027. This Actor provides a compatible POST /v1/audio/speech endpoint with independent English voices from $0.009 per started 1,000 input characters: change the base URL and API key, keep model, input, voice, response_format and speed, and remove nonempty instructions and SSE stream_format.

What changes on 2027-01-06

OpenAI's deprecations page, rechecked 3 October 2026, lists tts-1, tts-1-hd, gpt-4o-mini-tts-2025-03-20 and gpt-4o-mini-tts-2025-12-15 for removal on 6 January 2027 following its 1 October announcement. The recommended replacement, gpt-realtime-2.1-mini, uses the Realtime API and token-based audio pricing; that recommendation is not a guarantee that existing /audio/speech clients work unchanged. Plan and test the API migration before the removed models stop accepting requests.

Minimal migration diff

Change the base URL and API key. Keep model, input, voice, response_format and speed; remove nonempty instructions and SSE stream_format. Standby streams MP3, Opus, AAC and PCM up to 4,096 characters; WAV and FLAC are buffered and must be 500 characters or fewer. For example:

diff
  client = OpenAI(
-     api_key=os.environ["OPENAI_API_KEY"],
+     api_key=os.environ["APIFY_TOKEN"],
+     base_url="https://dropin-apis--tts-compat.apify.actor/v1",
  )
  with client.audio.speech.with_streaming_response.create(
      model="tts-1", input="Hello, world.", voice="amber", response_format="mp3"
  ) as response:
      response.stream_to_file("speech.mp3")

Or plain HTTP:

bash
curl --fail-with-body \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  --data '{"model":"tts-1","input":"Hello, world.","voice":"amber","response_format":"mp3"}' \
  'https://dropin-apis--tts-compat.apify.actor/v1/audio/speech' \
  --output speech.mp3

Confirm the real deployed hostname on the Actor's Endpoints tab before going live. model is accepted as a compatibility identifier only — tts-1, tts-1-hd and gpt-4o-mini-tts all select the same local model underneath, with no HD/quality tier. voice accepts the same parameter values (alloy, echo, nova, and so on) mapped to distinct independent local voices — re-test pronunciation and voice choice, since the acoustic output is different. Remove any nonempty instructions or SSE stream_format options: both return HTTP 400 here.

How it compares

OptionPriceNotes
This Actor (Standby)$0.009 per 1,000 charactersIndependent local Kokoro voices, English only, compatible request/response shape, no voice cloning
This Actor (Store/MCP batch run)$0.12 per 1,000 charactersSame voices; all six formats up to 4,096 characters, downloadable stored audio
OpenAI's recommended replacement (gpt-realtime-2.1-mini)$20 per 1M audio output tokensToken-priced realtime-voice-agent model, not a simple per-character TTS rate — needs its own cost modeling, not a straight swap
Kokoro Text to Speech on Apify$0.08 per 1,000 characters (per its own Store listing)Also Kokoro-based; 54 voices across 9 languages advertised, vs. this Actor's 13 English-only voices
Self-hosting Kokoro or Piper yourselfYour compute, no per-call feeNo managed endpoint, no OpenAI-compatible request shape out of the box, you run and scale it

This is an independent speech service, not an OpenAI product. Compatibility covers the request shape and the audio response format — not identical voices, pronunciation, prosody or model quality, and there is no affiliation with OpenAI or any other provider. Every successful response carries X-Voice-Provider and X-Local-Voice disclosure headers naming the actual local voice used. The voice parameter values above (alloy, echo, etc.) are accepted only as compatibility identifiers so existing request bodies keep working — they are not a claim that you get OpenAI's own voices.

Honest latency note

The live incremental-billing verification on 3 October 2026 streamed 4,096 characters of MP3 from Apify at 8 GB/two ONNX threads: first bytes arrived after 10.50 seconds, the body completed after 482.40 seconds, and all 264.70 seconds of audio decoded successfully. Five character events were acknowledged separately as the stream progressed. Historical measurements before incremental billing were 8.52/119.34 seconds (first byte/full body) for 1,000 characters and 7.31/503.00 seconds for 4,096. These proxy timings include a readiness GET; they are observations, not a latency or cold-start guarantee, and the difference is not a controlled speedup comparison. Streaming allows a long body to continue after the platform's five-minute first-response window, with a 270-second application first-byte deadline and a further 1,200-second stream deadline.

Use streamed MP3, Opus, AAC or PCM for up to 4,096 characters. HTTP WAV and FLAC are buffered and capped at 500 characters; longer input returns 413 use_batch_mode before synthesis or any character charge. A normal Store/MCP batch run supports all six formats up to 4,096 characters with a 1,200-second run timeout and returns stored audio. Generation runs on CPU and has no real-time guarantee; detailed measurements and conditions are in the Actor's VERIFY.md.

Pricing

Pay per event, no subscription. A completed Standby request charges characters-1k at $0.009 per started 1,000 Unicode characters (spaces count): 250 or 1,000 characters cost $0.009, 1,001 cost $0.018, 4,096 cost $0.045. Before synthesis, a full-input budget precheck rejects an insufficient cap with 402 and no work or audio. Streaming then charges one event before the first audio byte and each further event before releasing audio that can cross the next original-input 1,000-character boundary. Whole-input pronunciation validation and each new unit's first synthesis/encoding preflight happen before its charge; synthesis or encoding failure before a new unit starts does not charge that unit.

First-window validation, synthesis, encoding-preflight and deadline failures are free of character charges. If billing fails before audio starts, the response is 503 with no audio, or 402 when a cap is reached. A later synthesis, encoding or billing failure truncates the HTTP stream without appended JSON or a clean successful EOF: already-started units remain charged, later units do not. Clients must reject incomplete streams. A word spanning a billing boundary is kept whole and requires the next unit before its audio; trailing spaces count with the last audio window.

Normal Store/MCP batch-characters-1k remains $0.12 per started 1,000 characters, billed only after complete encoding; 4,096 characters cost $0.60. Apify's separate instance-start event can still apply to failed requests. Billing and network delivery are not atomic, so a connection loss after a successful charge can still prevent receipt of that window; there is no automatic refund or idempotency key, and a retry is a new chargeable request. X-Billed-Units is the full-input count for a completed response, not a final charge report for an interrupted stream; X-Billing-Mode: incremental identifies the streamed billing path.

Open Text-to-Speech API on Apify

FAQ

Is this the same model as OpenAI's tts-1?

No. It runs Kokoro-82M, an independent open local model, through quantized ONNX on CPU. The request and response shape are compatible so your existing client code keeps working; the voices, pronunciation and audio quality are not identical to OpenAI's.

What happens to my code after January 6, 2027?

Any request to OpenAI's API naming tts-1, tts-1-hd, gpt-4o-mini-tts-2025-03-20 or gpt-4o-mini-tts-2025-12-15 will fail once OpenAI removes them. Point base_url at this Actor, use an Apify token, keep the five supported speech fields and remove nonempty instructions and SSE options; re-test the independent voices and observe the format/input limits above.

Does this support Arabic or other languages?

No — English only in this version. An unsupported script returns HTTP 400 rather than silently mangled audio.

Why would a short, ordinary piece of text return an error?

The no-espeak English pronunciation lexicon doesn't cover every spelling, so an unrecognized name or acronym is rejected (HTTP 400) rather than silently mispronounced or dropped. Rephrase or spell it out.

Can an AI agent call this through Apify MCP?

Yes — a normal call-actor run synthesizes the text and returns one dataset row with a downloadable audio URL, billed at the batch rate ($0.12 per 1,000 characters). An HTTP-capable agent can call the cheaper Standby endpoint ($0.009) directly instead.

What audio formats are supported?

MP3 (default), Opus, AAC, FLAC, WAV and headerless PCM (signed 16-bit little-endian, mono, 24 kHz) — decode PCM with explicit parameters; it is not a WAV file.