
Tavus CVI
AI avatar model · Available
TavusConversational video with perception, so the avatar can see the person it is talking to.
Animates
Recruitment fit
90
Per hour
$60.00
Per minute
$1.00
30-min interview
$30.00
About
Tavus CVI is not one model but a stack: Phoenix renders the face, Raven watches the user through their webcam, and Sparrow decides when to speak. That perception layer is what separates it from a talking-head renderer, it can notice that a candidate is reading from a second screen, or that they stopped mid-sentence to think rather than because they were finished. For screening interviews that behaviour is the whole point, and it is the reason to pay a real premium over a pure lip-sync model.
- Conversational video interface
- Visual perception (Raven)
- Turn-taking model (Sparrow)
- Replica training from consent video
- Stock personas
- Real-time WebRTC
- 30+ languages
- Wired into HumanLike avatars: Yes, face_id provider
What it animates
- Face
- Full
- Lip-sync
- Full
- Upper torso
- Partial
- Hands
- None
- Full body
- None
- Prompt artifacts
- None
The Phoenix models generate the whole face, so brow, eyes and head motion move with the speech rather than only the mouth. Phoenix-4.5 adds natural torso and shoulder motion, and Phoenix-4 supports full body. Hands stay out of shot.
Source: Tavus documentation and HumanLike hands-on assessment
Strengths
- Reads the candidate’s camera
- Handles thinking pauses
- Recruiter replicas at scale
Watch out for
- 200× the cost of Homo 1
- Webcam consent and retention duties
- Replicas need written consent
Best for First-round conversational screening · Employer-brand videos personalised per candidate · Proctored assessments
3 variants

Phoenix-4.5Max
The newest face model. Usable in about a minute from a photo as a watermarked preview, then tuned in the background. Natural torso, shoulder and hair motion; chest-up or waist-up framing.

Phoenix-4High
The full-body option. Keeps complex clothing texture exact and has the highest fidelity from video, but is only usable once the whole training job finishes.

Phoenix-3 (legacy)Balanced
The previous model. Video training only; rarely the right choice over Phoenix-4.5 or Phoenix-4.
The numbers
Published benchmarks
- Lip-sync accuracy
- 90%Estimate
- Utterance-to-utterance latency
- 1000 msEstimate
Published by us as an approximation (~90%).
Published by us as an approximation (~1.0s).
Paper
Tavus CVI: a conversation stack, not just a renderer
Tavus CVI joins a face renderer, a perception model and a turn-taking model behind one real-time video interface. The renderer makes it look human; the other two make the conversation behave like one.
Read the paper