
GPT Transcribe
audio model · Available
OpenAIThe only model in our harness that does not degrade under background noise.
Fit and price
Recruitment fit
87
Per hour
$0.270
Per minute
$0.0045
30-min interview
$0.135
About
GPT Transcribe holds the "Smart" tier in HumanLike Coworker, and the reason is counter-intuitive. On clean audio it is the *worst* of the four models we benchmarked, 0.29 WER against mini’s 0.12. Add background noise and the ordering inverts: it lands at 0.26 while mini collapses to 0.42. It is the only model we tested that does not get worse when the room does, and real dictation happens in real rooms.
- Noise-robust
- Technical vocabulary
- Proper nouns
- HTTP file endpoint
Strengths and watch-outs
Strengths
- Holds accuracy under noise
- Handles jargon and proper nouns
- $0.0045/min, transcribe every call
Watch out for
- Balanced beats Smart on clean audio
- Small samples, validate yourself
Best for Interview recordings · Screening call notes · Technical role discussions
3 variants

SmartHigh
Highest quality for complex or technical content, and the most noise-tolerant.
1.4s
latency
1×
rel. cost
86
quality

BalancedBalanced
Best clean-audio accuracy in our harness, and the worst under noise.
0.5×
rel. cost
78
quality

Live (streaming)Fast
WebSocket-only. Text streams while the speaker is still talking, ~300ms of text remains after they stop.
300ms
latency
3.8×
rel. cost
80
quality
The numbers
HumanLike evals
- Word error rate (clean audio)
- 0.29Our eval
- Word error rate (with additive noise)
- 0.26Our eval
- Median latency (9.2s clip)
- 1390 msOur eval
Small sample. Indicative, not definitive. Worst of four models on clean audio.
Best of four. The only model that did not degrade under noise.
Equivalent to gpt-4o-transcribe (1.45s), not an improvement.