audio model

The best Greek in our harness, and the only model here with speaker attribution.
Two models share this name and they are not interchangeable. The Live sibling is the fastest streaming transcription we have measured — 666ms of text still arriving after an English speaker stops, against 929ms for OpenAI’s streaming model. The recorded-audio model is slower but carries speaker attribution and word-level timestamps, which nothing else in our stack offers. For a panel interview that distinction decides which one you want.
Variants
Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.
gemini-3.5-transcribe-liveWebSocket streaming. Fastest tail latency we have measured — 666ms EN / 876ms EL.
Live captions · Real-time interview notes
gemini-3.5-transcribeSlower (2.4s EN / 2.6s EL end-to-end on opus) but adds speaker attribution and word timestamps.
Panel interviews · Anything needing who-said-what
Fit
The same model is a different proposition depending on what you point it at.
Recruitment fit
The pick for panel interviews and any non-English hiring — speaker attribution and Greek accuracy are both unmatched here.
Best for
Strengths
Watch out for
The numbers
HumanLike evals
Run on our own harness.
Best Greek of any model in our harness except gpt-transcribe.
Against 929ms for gpt-live-transcribe and 1197ms for the Groq default.
Compare