Side by side
Up to 4 models, normalised onto the same cost basis and scored through the lens you pick.
| Cost of talking | ||
| Per minute | $1.00Estimate | Not conversational |
| Per hour | $60.00 | — |
| 30-min interview | $30.00 | — |
| List price | $1/minEstimate | Plan-basedEstimate |
| Recruitment fit | ||
| Overall | 90/100 | 55/100 |
| Verdict | The most complete option for conversational screening — the perception layer is worth the premium when a real person is on the other end. | Excellent for onboarding and policy video; structurally unable to do live screening. |
| Best for | First-round conversational screeningEmployer-brand videos personalised per candidateProctored assessments where attention signals matter | Onboarding modulesCompliance trainingCareers-page brand film |
| Watch out for |
|
|
| What it animates | ||
| Coverage | ||
| Face | Full | Full |
| Lip-sync | Full | Full |
| Upper torso | Partial | Full |
| Hands | None | Partial |
| Full body | None | Full |
| Capabilities | ||
| Modalities | avatar, video, speech | avatar, video |
| Context window | — | — |
| Variants | Phoenix-3 (render), Raven-0 (perception), Sparrow-0 (turn-taking), Hummingbird-0 (lip-sync) | Studio |
| On HumanLike | AI Personas, Interview | Not yet |
| HumanLike evals | ||
| Wired into HumanLike avatars | Yes — face_id providerOur eval | — |
| Published benchmarks | ||
| Lip-sync accuracy | 90%Estimate | 94%Estimate |
| Utterance-to-utterance latency | 1000 msEstimate | minutesEstimate |
Cost per hour assumes 150 wpm and a 40% AI speaking share. Hover a figure for its full derivation.