AI avatar model

Tavus logo

Tavus CVI

Tavus

Conversational video with perception — the avatar can see the person it is talking to.

Tavus CVI is not one model but a stack: Phoenix renders the face, Raven watches the user through their webcam, and Sparrow decides when to speak. That perception layer is what separates it from a talking-head renderer — it can notice that a candidate is reading from a second screen, or that they stopped mid-sentence to think rather than because they were finished. For screening interviews that behaviour is the whole point, and it is the reason to pay a real premium over a pure lip-sync model.

Use it on HumanLike
CapabilitiesConversational video interfaceVisual perception (Raven)Turn-taking model (Sparrow)Replica training from consent videoStock personasReal-time WebRTC30+ languages

Sample output

What it actually produces.

One clip is worth more than any benchmark row. Where we have not published a sample yet, the frame below holds the space it will occupy.

Demo clip

No sample of Tavus CVI uploaded yet.

Phoenix-3 replica answering a screening question.

What it animates

Head to toe.

A model that only drives the mouth looks wrong the moment the other person starts talking — nothing on screen moves. Filled means driven, outlined means limited or looped, greyed means static.

Face
Full
Lip-sync
Full
Upper torso
Partial
Hands
None
Full body
None

Phoenix-3 does full-face reenactment, so brow, eyes and head motion move with the speech rather than only the mouth. Framing is head-and-shoulders: shoulders drift naturally but hands are out of shot.

Source: HumanLike hands-on assessment of Tavus CVI output

Variants

4 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Phoenix-3 (render)

Max
phoenix-3

The rendering model. Full-face reenactment rather than mouth-only compositing, so expression and micro-head-motion move with the speech.

Candidate-facing screening calls · High-trust brand video

1.0s
latency
rel. cost
90
quality

Raven-0 (perception)

High
raven-0

Visual perception. Reads the user’s camera for attention, emotion and environment — this is what makes the conversation feel two-way.

Proctored screening · Engagement signals

rel. cost
86
quality

Sparrow-0 (turn-taking)

Balanced
sparrow-0

Decides when to speak. Handles pauses, interruptions and thinking silences without talking over the candidate.

Natural interview pacing

600ms
latency
1.5×
rel. cost
84
quality

Hummingbird-0 (lip-sync)

Fast
hummingbird-0

Lip-sync only, applied to existing footage. The cheap path when you do not need conversation.

Localising recorded video · Bulk personalised outreach

rel. cost
78
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

90/100

Recruitment fit

The most complete option for conversational screening — the perception layer is worth the premium when a real person is on the other end.

Best for

  • First-round conversational screening
  • Employer-brand videos personalised per candidate
  • Proctored assessments where attention signals matter

Strengths

  • Reads the candidate’s camera, so the interview reacts rather than just plays
  • Turn-taking handles thinking pauses without cutting candidates off
  • Replica training lets a real recruiter’s face front the screen at scale

Watch out for

  • At roughly $1/min it is 200× the cost of Homo 1 — model your per-interview cost before committing
  • Recording a candidate’s webcam raises consent and data-retention obligations; several jurisdictions require explicit notice for automated video assessment
  • A trained replica of a real employee needs that employee’s written consent

The numbers

Measured, not asserted.

HumanLike evals

Run on our own harness.

Wired into HumanLike avatars
Yes — face_id providerOur eval

Selectable as a video provider on any AI persona; requires a trained Tavus face_id.

Published benchmarks

Vendor and independent figures.

Lip-sync accuracy
90%Estimate

Published by us as an approximation (~90%).

Utterance-to-utterance latency
1000 msEstimate

Published by us as an approximation (~1.0s).