AI avatar model

HumanLike logo

HumanLike Homo 1

HumanLike

$0.005 a minute, 300ms end to end. The price disruption is the product.

Homo 1 is our own real-time avatar. It does one thing — put a convincing human face on a live conversation — and it does it at a price that changes what you can build. At $0.005/min an always-on avatar costs less per hour than the electricity to display it, which makes it viable for surfaces nobody would put a $1/min avatar on: every job posting, every candidate FAQ, every product page.

Use it on HumanLike
CapabilitiesReal-time WebRTCSub-second latencyLip-sync from any TTSCustom faces

Sample output

What it actually produces.

One clip is worth more than any benchmark row. Where we have not published a sample yet, the frame below holds the space it will occupy.

Demo clip

No sample of HumanLike Homo 1 uploaded yet.

Homo 1 holding a live candidate conversation at 300ms.

What it animates

Head to toe.

A model that only drives the mouth looks wrong the moment the other person starts talking — nothing on screen moves. Filled means driven, outlined means limited or looped, greyed means static.

Face
Full
Lip-sync
Full
Upper torso
Partial
Hands
None
Full body
None

Face and lips are fully driven at 99% lip-sync. Torso carries a subtle idle motion so the frame never freezes between turns; hands and full body are out of frame.

Source: HumanLike Homo 1 rendering pipeline

Variants

1 way to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Real-time

Balanced
homo-1-realtime

Live conversational rendering over WebRTC at 300ms end-to-end.

Live candidate chat · Always-on site avatars

300ms
latency
rel. cost
92
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

88/100

Recruitment fit

The cost floor for candidate-facing video — cheap enough to leave running on every job page, not just the interview.

Best for

  • Job-page avatars
  • Candidate FAQ and chat
  • High-volume first contact

Strengths

  • Two orders of magnitude cheaper per minute than CVI-class providers
  • 300ms latency reads as conversational rather than laggy
  • Native to the platform: no third-party data processor in the candidate path

Watch out for

  • No visual perception layer — it does not see the candidate the way Tavus Raven does
  • Proctoring and attention signals need a separate component

The numbers

Measured, not asserted.

Published benchmarks

Vendor and independent figures.

Lip-sync accuracy
99%Our eval
End-to-end latency
300 msOur eval