The model index

Every model,
priced by the hour
it actually talks.

Text, audio, speech, avatars and video are billed five different ways. We convert all of them to one number — cost per minute and per hour of live conversation — then score each model twice, once for hiring and once for general work.

Every figure carries its source. Nothing appears as a bare claim.

19
Models
5
Modalities
10
On HumanLike
No. 01Top of the recruitment index
OpenAI logo

GPT-6

OpenAI

The strongest general model for hiring work — but only worth its price on the judgement-heavy steps, not on parsing.
Read the full profile
93/100
Recruitment fit
$0.698
Per hour talking
$0.012
Per minute
$0.349
30-min interview
Best for
Structured interview scoring against a rubricComparing a shortlist against a job specDrafting candidate feedback that a human will edit
ModelRecruit. fitCost
02Anthropic logo
Claude Opus 5
Anthropic/Available

Long-horizon reasoning and the steadiest instruction-following in the catalogue.

91/100
$5.76/hr
$0.096/min
03Tavus logo
Tavus CVI
Tavus/Available

Conversational video with perception — the avatar can see the person it is talking to.

90/100
$60.00/hr
$1.00/min
04Google logo
Gemini 3.5 Transcribe
Google/Available

The best Greek in our harness, and the only model here with speaker attribution.

89/100
05HumanLike logo
HumanLike Homo 1
HumanLike/Available

$0.005 a minute, 300ms end to end. The price disruption is the product.

88/100
$0.300/hr
$0.0050/min
06OpenAI logo
GPT Transcribe
OpenAI/Available

The only model in our harness that does not degrade under background noise.

87/100
$0.270/hr
$0.0045/min
07OpenAI logo
GPT-5.6
OpenAI/Available

The model behind HumanLike Coworker — three tiers, one API surface.

86/100
$0.499/hr
$0.0083/min
08Google logo
Gemini 3.5
Google/Available

Huge context and native audio/video understanding at aggressive prices.

84/100
$0.499/hr
$0.0083/min
09ElevenLabs logo
ElevenLabs v3
ElevenLabs/Available

The most expressive synthetic voice available, and the best voice cloning.

84/100
$3.69/hr
$0.062/min
10OpenAI logo
GPT-5
OpenAI/Superseded

The model that introduced the reasoning-effort dial.

78/100
$0.499/hr
$0.0083/min
11OpenAI logo
OpenAI TTS
OpenAI/Available

Steerable voice — you describe the delivery in a prompt rather than tuning sliders.

76/100
$0.308/hr
$0.0051/min
12HeyGen logo
HeyGen Interactive Avatar
HeyGen/Available

The most polished stock-avatar library, with real-time streaming on top.

72/100
$30.00/hr
$0.500/min
13OpenAI logo
GPT-4o
OpenAI/Superseded

The first natively multimodal GPT — text, vision and audio in one model.

66/100
$0.949/hr
$0.016/min
14ByteDance logo
Seedance 2.5
ByteDance/Available

Text-to-video with natively synchronised audio — one pass, not a dub.

64/100
$432/hr
$7.20/min
15Groq logo
Whisper v3 Turbo on Groq
Groq/Available

The fast tier — cheapest path to a good-enough transcript.

62/100
16OpenAI logo
Sora 2
OpenAI/Available

The most physically coherent generated video — objects keep obeying physics across a shot.

58/100
$360/hr
$6.00/min
17Synthesia logo
Synthesia
Synthesia/Available

Pre-rendered only — the enterprise standard for produced training video.

55/100
18OpenAI logo
GPT-4
OpenAI/Superseded

The generational jump that made LLMs viable for real work.

48/100
$3.75/hr
$0.062/min
19OpenAI logo
GPT-3.5 Turbo
OpenAI/Retired

The model that started it all — and the cost floor everything is measured against.

22/100
$0.187/hr
$0.0031/min

Methodology

Where these numbers come from.

Published
Figures the vendor or an independent benchmark published. The citation is on every number.
Our eval
Measured on our own harness, with the trial count and clip length recorded. Small samples — indicative, not definitive.
Estimate
Not yet verified — usually an unreleased model or plan-based pricing. A placeholder, not a quote.

Cost per hour assumes a 150 wpm speaking pace and a 40% AI speaking share — the shape of a screening interview, where the model asks and the candidate answers. For token-billed models it also assumes three turns a minute against a cached prompt prefix. Hover any figure to see its full derivation.

Run any of these on HumanLike.

One integration for text, voice, avatars and video — swap models without rewriting the pipeline.