
GPT-4o
text model · Superseded
OpenAIThe first natively multimodal GPT, text, vision and audio in one model.
Fit and price
Recruitment fit
66
Per hour
$0.949
Per minute
$0.016
30-min interview
$0.474
About
GPT-4o collapsed three separate pipelines (transcribe → reason → speak) into a single model, which is why it still turns up in real-time voice stacks. For text-only work it has been comprehensively superseded, but the realtime variant remains a reasonable low-latency voice option.
- Vision
- Audio in/out
- Tool use
- Structured outputs
Strengths and watch-outs
Strengths
- Low-latency speech-to-speech
- Very cheap Mini for bulk
Watch out for
- 128K context is limiting
- Weak reasoning
Best for Latency-critical voice screening
3 variants

StandardBalanced
The full multimodal model.
16×
rel. cost
74
quality

MiniMini
Cheap, fast, and still multimodal.
1×
rel. cost
61
quality

RealtimeFast
Speech-to-speech over WebSocket, no transcription hop.
320ms
latency
40×
rel. cost
70
quality
The numbers
Published benchmarks
- MMLU
- 88.7%Published