Side by side
Up to 4 models, normalised onto the same cost basis and scored through the lens you pick.
| Cost of talking | ||
| Per minute | $0.500Estimate | $1.00Estimate |
| Per hour | $30.00 | $60.00 |
| 30-min interview | $15.00 | $30.00 |
| List price | $0.5/minEstimate | $1/minEstimate |
| Recruitment fit | ||
| Overall | 72/100 | 90/100 |
| Verdict | Good for employer-brand content, weaker as a live interviewer than Tavus. | The most complete option for conversational screening — the perception layer is worth the premium when a real person is on the other end. |
| Best for | Employer-brand videoMultilingual job adsOnboarding content | First-round conversational screeningEmployer-brand videos personalised per candidateProctored assessments where attention signals matter |
| Watch out for |
|
|
| What it animates | ||
| Coverage | ||
| Face | Full | Full |
| Lip-sync | Full | Full |
| Upper torso | Full | Partial |
| Hands | Partial | None |
| Full body | Partial | None |
| Capabilities | ||
| Modalities | avatar, video | avatar, video, speech |
| Context window | — | — |
| Variants | Interactive, Studio (rendered) | Phoenix-3 (render), Raven-0 (perception), Sparrow-0 (turn-taking), Hummingbird-0 (lip-sync) |
| On HumanLike | Not yet | AI Personas, Interview |
| HumanLike evals | ||
| Wired into HumanLike avatars | — | Yes — face_id providerOur eval |
| Published benchmarks | ||
| Lip-sync accuracy | 92%Estimate | 90%Estimate |
| Utterance-to-utterance latency | 1200 msEstimate | 1000 msEstimate |
Cost per hour assumes 150 wpm and a 40% AI speaking share. Hover a figure for its full derivation.