speech model

Steerable voice — you describe the delivery in a prompt rather than tuning sliders.
OpenAI’s TTS models are less expressive than ElevenLabs at the top end but introduce a genuinely useful idea: you instruct the delivery in natural language ("warm, unhurried, like you are explaining to a nervous candidate"). For screening flows where tone should shift with context, that is easier to control than a fixed voice preset.
Variants
Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.
gpt-4o-mini-ttsSteerable delivery at low cost.
tts-1-hdHigher-fidelity preset voices, no steering.
Fit
The same model is a different proposition depending on what you point it at.
Recruitment fit
The cost-effective choice when tone needs to vary by stage but the voice does not need to be anyone real.
Best for
Strengths
Watch out for