audio model

The only model in our harness that does not degrade under background noise.
GPT Transcribe holds the "Smart" tier in HumanLike Coworker, and the reason is counter-intuitive. On clean audio it is the *worst* of the four models we benchmarked — 0.29 WER against mini’s 0.12. Add background noise and the ordering inverts: it lands at 0.26 while mini collapses to 0.42. It is the only model we tested that does not get worse when the room does, and real dictation happens in real rooms.
Variants
Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.
gpt-transcribeHighest quality for complex or technical content, and the most noise-tolerant.
Interviews recorded in real rooms · Technical content
gpt-4o-mini-transcribeBest clean-audio accuracy in our harness — and the worst under noise.
Studio-quality audio · Cost-sensitive volume
gpt-live-transcribeWebSocket-only. Text streams while the speaker is still talking — ~300ms of text remains after they stop.
Live captioning · Real-time interview notes
Fit
The same model is a different proposition depending on what you point it at.
Recruitment fit
The right transcription default for interviews, because interviews are recorded in noisy rooms on laptop mics.
Best for
Strengths
Watch out for
The numbers
HumanLike evals
Run on our own harness.
Small sample. Indicative, not definitive. Worst of four models on clean audio.
Best of four. The only model that did not degrade under noise.
Equivalent to gpt-4o-transcribe (1.45s), not an improvement.
Compare