Humanlike AI
All models
Google logo

Gemini 3.5 Transcribe

audio model · Available

Google

The best Greek in our harness, and the only model here with speaker attribution.

Gemini 3.5 TranscribeIdle · sample soon

Fit and price

Recruitment fit

89

Per hour

n/a

Per minute

n/a

30-min interview

n/a

About

Two models share this name and they are not interchangeable. The Live sibling is the fastest streaming transcription we have measured, 666ms of text still arriving after an English speaker stops, against 929ms for OpenAI’s streaming model. The recorded-audio model is slower but carries speaker attribution and word-level timestamps, which nothing else in our stack offers. For a panel interview that distinction decides which one you want.

  • Speaker attribution
  • Word-level timestamps
  • Streaming sibling
  • Strong multilingual

Strengths and watch-outs

Strengths

  • Speaker attribution for panels
  • Word-level timestamps for deep links
  • Best Greek WER (0.03)

Watch out for

  • Live and Recorded not interchangeable
  • Needs a billed key

Best for Panel interviews · Multilingual hiring · Evidence-linked interview notes

2 variants

  • Google logo

    LiveFast

    WebSocket streaming. Fastest tail latency we have measured, 666ms EN / 876ms EL.

    666ms

    latency

    1×

    rel. cost

    94

    quality

  • Google logo

    RecordedHigh

    Slower (2.4s EN / 2.6s EL end-to-end on opus) but adds speaker attribution and word timestamps.

    2.4s

    latency

    1×

    rel. cost

    90

    quality

The numbers

HumanLike evals

Word error rate (English)
0.06Our eval
Word error rate (Greek)
0.03Our eval

Best Greek of any model in our harness except gpt-transcribe.

Text still arriving after speaker stops (EN)
666 msOur eval

Against 929ms for gpt-live-transcribe and 1197ms for the Groq default.

Compare with