speech model

OpenAI logo

OpenAI TTS

OpenAI

Steerable voice — you describe the delivery in a prompt rather than tuning sliders.

OpenAI’s TTS models are less expressive than ElevenLabs at the top end but introduce a genuinely useful idea: you instruct the delivery in natural language ("warm, unhurried, like you are explaining to a nervous candidate"). For screening flows where tone should shift with context, that is easier to control than a fixed voice preset.

Use it on HumanLike
CapabilitiesPrompted deliveryStreamingLow costPreset voices

Variants

2 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Mini TTS

Balanced
gpt-4o-mini-tts

Steerable delivery at low cost.

400ms
latency
rel. cost
82
quality

TTS-1 HD

High
tts-1-hd

Higher-fidelity preset voices, no steering.

700ms
latency
rel. cost
78
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

76/100

Recruitment fit

The cost-effective choice when tone needs to vary by stage but the voice does not need to be anyone real.

Best for

  • Automated candidate updates
  • Interview instructions
  • High-volume notifications

Strengths

  • Prompted delivery adapts tone per stage without swapping voices
  • Materially cheaper than ElevenLabs at volume
  • No cloning means no consent problem

Watch out for

  • Less expressive at the top end
  • No custom voice cloning

Compare

Worth putting side by side.