audio model

OpenAI logo

GPT Transcribe

OpenAI

The only model in our harness that does not degrade under background noise.

GPT Transcribe holds the "Smart" tier in HumanLike Coworker, and the reason is counter-intuitive. On clean audio it is the *worst* of the four models we benchmarked — 0.29 WER against mini’s 0.12. Add background noise and the ordering inverts: it lands at 0.26 while mini collapses to 0.42. It is the only model we tested that does not get worse when the room does, and real dictation happens in real rooms.

Use it on HumanLike
CapabilitiesNoise-robustTechnical vocabularyProper nounsHTTP file endpoint

Variants

3 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Smart

High
gpt-transcribe

Highest quality for complex or technical content, and the most noise-tolerant.

Interviews recorded in real rooms · Technical content

1.4s
latency
rel. cost
86
quality

Balanced

Balanced
gpt-4o-mini-transcribe

Best clean-audio accuracy in our harness — and the worst under noise.

Studio-quality audio · Cost-sensitive volume

0.5×
rel. cost
78
quality

Live (streaming)

Fast
gpt-live-transcribe

WebSocket-only. Text streams while the speaker is still talking — ~300ms of text remains after they stop.

Live captioning · Real-time interview notes

300ms
latency
3.8×
rel. cost
80
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

87/100

Recruitment fit

The right transcription default for interviews, because interviews are recorded in noisy rooms on laptop mics.

Best for

  • Interview recordings
  • Screening call notes
  • Technical role discussions

Strengths

  • Only model in our harness that holds accuracy under background noise
  • Handles technical vocabulary and proper nouns — job titles, tool names, company names
  • Cheap enough at $0.0045/min to transcribe every call

Watch out for

  • On clean studio audio the Balanced tier is genuinely more accurate — do not blanket-apply Smart
  • Small benchmark samples: re-validate on your own audio before standardising

The numbers

Measured, not asserted.

HumanLike evals

Run on our own harness.

Word error rate (clean audio)
0.29Our eval

Small sample. Indicative, not definitive. Worst of four models on clean audio.

Word error rate (with additive noise)
0.26Our eval

Best of four. The only model that did not degrade under noise.

Median latency (9.2s clip)
1390 msOur eval

Equivalent to gpt-4o-transcribe (1.45s), not an improvement.