text model

OpenAI logo

GPT-4o

OpenAI

The first natively multimodal GPT — text, vision and audio in one model.

GPT-4o collapsed three separate pipelines (transcribe → reason → speak) into a single model, which is why it still turns up in real-time voice stacks. For text-only work it has been comprehensively superseded, but the realtime variant remains a reasonable low-latency voice option.

CapabilitiesVisionAudio in/outTool useStructured outputs

Variants

3 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Standard

Balanced
gpt-4o

The full multimodal model.

16×
rel. cost
74
quality

Mini

Mini
gpt-4o-mini

Cheap, fast, and still multimodal.

rel. cost
61
quality

Realtime

Fast
gpt-4o-realtime

Speech-to-speech over WebSocket, no transcription hop.

Live voice screening calls

320ms
latency
40×
rel. cost
70
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

66/100

Recruitment fit

Only worth choosing today for its realtime voice variant; the text tiers are outclassed.

Best for

  • Live voice screening where latency dominates

Strengths

  • Realtime speech-to-speech is genuinely low latency
  • Mini is very cheap for bulk parsing

Watch out for

  • 128K context is limiting for long interview transcripts
  • Reasoning is well behind current models

The numbers

Measured, not asserted.

Published benchmarks

Vendor and independent figures.

MMLU
88.7%Published

Compare

Worth putting side by side.