audio model

Google logo

Gemini 3.5 Transcribe

Google

The best Greek in our harness, and the only model here with speaker attribution.

Two models share this name and they are not interchangeable. The Live sibling is the fastest streaming transcription we have measured — 666ms of text still arriving after an English speaker stops, against 929ms for OpenAI’s streaming model. The recorded-audio model is slower but carries speaker attribution and word-level timestamps, which nothing else in our stack offers. For a panel interview that distinction decides which one you want.

Use it on HumanLike
CapabilitiesSpeaker attributionWord-level timestampsStreaming siblingStrong multilingual

Variants

2 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Live

Fast
gemini-3.5-transcribe-live

WebSocket streaming. Fastest tail latency we have measured — 666ms EN / 876ms EL.

Live captions · Real-time interview notes

666ms
latency
rel. cost
94
quality

Recorded

High
gemini-3.5-transcribe

Slower (2.4s EN / 2.6s EL end-to-end on opus) but adds speaker attribution and word timestamps.

Panel interviews · Anything needing who-said-what

2.4s
latency
rel. cost
90
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

89/100

Recruitment fit

The pick for panel interviews and any non-English hiring — speaker attribution and Greek accuracy are both unmatched here.

Best for

  • Panel interviews
  • Multilingual hiring
  • Evidence-linked interview notes

Strengths

  • Speaker attribution turns a panel recording into an attributable transcript
  • Word-level timestamps let you deep-link a claim back to the moment it was said
  • Best measured Greek WER in our harness (0.03)

Watch out for

  • Live and Recorded are different model codes — you cannot swap one for the other
  • Requires a billed key; the free tier 429s under sustained use

The numbers

Measured, not asserted.

HumanLike evals

Run on our own harness.

Word error rate (English)
0.06Our eval
Word error rate (Greek)
0.03Our eval

Best Greek of any model in our harness except gpt-transcribe.

Text still arriving after speaker stops (EN)
666 msOur eval

Against 929ms for gpt-live-transcribe and 1197ms for the Groq default.