text model

The first natively multimodal GPT — text, vision and audio in one model.
GPT-4o collapsed three separate pipelines (transcribe → reason → speak) into a single model, which is why it still turns up in real-time voice stacks. For text-only work it has been comprehensively superseded, but the realtime variant remains a reasonable low-latency voice option.
Variants
Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.
gpt-4oThe full multimodal model.
gpt-4o-miniCheap, fast, and still multimodal.
gpt-4o-realtimeSpeech-to-speech over WebSocket, no transcription hop.
Live voice screening calls
Fit
The same model is a different proposition depending on what you point it at.
Recruitment fit
Only worth choosing today for its realtime voice variant; the text tiers are outclassed.
Best for
Strengths
Watch out for
The numbers
Published benchmarks
Vendor and independent figures.