Model paper
Tavus CVI: a conversation stack, not just a renderer
Tavus CVI joins a face renderer, a perception model and a turn-taking model behind one real-time video interface. The renderer makes it look human; the other two make the conversation behave like one.
HumanLike benchmark · Reviewed 2026-10-03 · Summary of public documentation
Architecture
- 01
Listen
The person talks over WebRTC, with their camera on.
- 02
Perceive (Raven)
Reads the camera for attention, emotion and surroundings, and passes that on as context.
- 03
Take turns (Sparrow)
Decides when the person has finished, handling pauses and interruptions.
- 04
Reply
A language model writes the answer and a voice speaks it.
- 05
Render (Phoenix)
A generated face, with expression and head motion, streams back as live video.
Why it is a stack
Most avatar products are renderers: you give them audio and they give you a talking face. Tavus CVI adds two models around that renderer. One watches the person, and one decides when to speak. Together they cover the two things a lip-sync model cannot do, which are noticing the other person and reacting to them.
Rendering
The face is generated, not a mouth pasted onto a still frame: brow, eyes and head motion move with the speech, and the face shows emotion while speaking and listening. Tavus offers three Phoenix models. Phoenix-4.5 is the newest: it adds more natural torso, shoulder, hair and clothing motion, and you can use a zero-shot preview within minutes while it tunes in the background (the preview carries a visible watermark). Phoenix-4 is the full-body option and keeps complex clothing texture exact, but is only usable once the whole training job finishes. Phoenix-3 is legacy, video-only, and rarely the right pick.
A face is trained from a short video, or from a single photo on Phoenix-4 and 4.5, and you can also pick a stock face. A face of a real employee needs that employee's written consent.
Perception
Raven reads the person's camera while they talk. For screening interviews this is the point of the product: it can notice that a candidate looks away to read from a second screen, or that they paused to think rather than because they were finished.
Recording a candidate's webcam raises consent and data-retention duties, and several jurisdictions require explicit notice for automated video assessment. Plan that before you switch it on.
Latency and cost
Our catalogue lists roughly one second from the end of an utterance to the start of the reply, and about $1 a minute. Both figures are estimates taken from Tavus's own comparison and have not been verified by us.
At that price an interview costs about 200 times as much as a lip-sync-only avatar such as HumanLike Homo 1, so the perception layer has to be worth it for your use.