text model

Google logo

Gemini 3.5

Google

Huge context and native audio/video understanding at aggressive prices.

Gemini 3.5 is the value pick when the input is large or non-textual. Its native audio and video understanding is why the same family powers our transcription tier — you can hand it a raw interview recording rather than a transcript.

Use it on HumanLike
Capabilities1M contextNative audioNative videoTool useStructured outputs

Variants

3 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Pro

High
gemini-3.5-pro

Frontier tier with the full 1M context.

rel. cost
91
quality

Flash

Balanced
gemini-3.5-flash

The price/performance sweet spot.

1.5×
rel. cost
82
quality

Flash-Lite

Mini
gemini-3.5-flash-lite

Cheapest tier for mechanical work.

rel. cost
68
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

84/100

Recruitment fit

The cheapest credible way to process raw interview recordings without a separate transcription step.

Best for

  • Interview recording analysis
  • Multilingual pipelines
  • Bulk CV parsing

Strengths

  • Accepts audio and video directly — no transcribe-then-reason hop
  • 1M context swallows an entire hiring round at once
  • Best measured Greek accuracy of anything we have benchmarked

Watch out for

  • Reasoning trails Opus 5 and GPT-6 on rubric consistency

The numbers

Measured, not asserted.

HumanLike evals

Run on our own harness.

Greek transcription WER (Live sibling)
0.03Our eval

Measured on gemini-3.5-transcribe-live, the audio sibling of this family.

Compare

Worth putting side by side.