AI avatar model

Conversational video with perception — the avatar can see the person it is talking to.
Tavus CVI is not one model but a stack: Phoenix renders the face, Raven watches the user through their webcam, and Sparrow decides when to speak. That perception layer is what separates it from a talking-head renderer — it can notice that a candidate is reading from a second screen, or that they stopped mid-sentence to think rather than because they were finished. For screening interviews that behaviour is the whole point, and it is the reason to pay a real premium over a pure lip-sync model.
Sample output
One clip is worth more than any benchmark row. Where we have not published a sample yet, the frame below holds the space it will occupy.
Demo clip
No sample of Tavus CVI uploaded yet.
What it animates
A model that only drives the mouth looks wrong the moment the other person starts talking — nothing on screen moves. Filled means driven, outlined means limited or looped, greyed means static.
Phoenix-3 does full-face reenactment, so brow, eyes and head motion move with the speech rather than only the mouth. Framing is head-and-shoulders: shoulders drift naturally but hands are out of shot.
Source: HumanLike hands-on assessment of Tavus CVI output
Variants
Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.
phoenix-3The rendering model. Full-face reenactment rather than mouth-only compositing, so expression and micro-head-motion move with the speech.
Candidate-facing screening calls · High-trust brand video
raven-0Visual perception. Reads the user’s camera for attention, emotion and environment — this is what makes the conversation feel two-way.
Proctored screening · Engagement signals
sparrow-0Decides when to speak. Handles pauses, interruptions and thinking silences without talking over the candidate.
Natural interview pacing
hummingbird-0Lip-sync only, applied to existing footage. The cheap path when you do not need conversation.
Localising recorded video · Bulk personalised outreach
Fit
The same model is a different proposition depending on what you point it at.
Recruitment fit
The most complete option for conversational screening — the perception layer is worth the premium when a real person is on the other end.
Best for
Strengths
Watch out for
The numbers
HumanLike evals
Run on our own harness.
Selectable as a video provider on any AI persona; requires a trained Tavus face_id.
Published benchmarks
Vendor and independent figures.
Published by us as an approximation (~90%).
Published by us as an approximation (~1.0s).
Compare