
Gemini 3.5 Transcribe
audio model · Available
GoogleThe best Greek in our harness, and the only model here with speaker attribution.
Fit and price
Recruitment fit
89
Per hour
n/a
Per minute
n/a
30-min interview
n/a
About
Two models share this name and they are not interchangeable. The Live sibling is the fastest streaming transcription we have measured, 666ms of text still arriving after an English speaker stops, against 929ms for OpenAI’s streaming model. The recorded-audio model is slower but carries speaker attribution and word-level timestamps, which nothing else in our stack offers. For a panel interview that distinction decides which one you want.
- Speaker attribution
- Word-level timestamps
- Streaming sibling
- Strong multilingual
Strengths and watch-outs
Strengths
- Speaker attribution for panels
- Word-level timestamps for deep links
- Best Greek WER (0.03)
Watch out for
- Live and Recorded not interchangeable
- Needs a billed key
Best for Panel interviews · Multilingual hiring · Evidence-linked interview notes
2 variants

LiveFast
WebSocket streaming. Fastest tail latency we have measured, 666ms EN / 876ms EL.
666ms
latency
1×
rel. cost
94
quality

RecordedHigh
Slower (2.4s EN / 2.6s EL end-to-end on opus) but adds speaker attribution and word timestamps.
2.4s
latency
1×
rel. cost
90
quality
The numbers
HumanLike evals
- Word error rate (English)
- 0.06Our eval
- Word error rate (Greek)
- 0.03Our eval
- Text still arriving after speaker stops (EN)
- 666 msOur eval
Best Greek of any model in our harness except gpt-transcribe.
Against 929ms for gpt-live-transcribe and 1197ms for the Groq default.