Model paper
Spatius: a 3D avatar rendered on the viewer's device
Spatius sends motion, not video. A server turns speech into synchronised face motion, and the browser SDK draws a 3D Gaussian splat avatar from it on the viewer's own GPU.
HumanLike benchmark · Reviewed 2026-10-03 · Summary of public documentation
Architecture
- 01
Speech
A text-to-voice service speaks the agent's reply.
- 02
Motion server
Spatius turns that speech into face and upper-body motion, aligned to the audio.
- 03
LiveKit
Audio and motion travel together over a LiveKit room.
- 04
Browser SDK
The Spatius SDK loads the avatar and draws it from the motion stream.
- 05
Viewer
The avatar is rendered locally, so it can be framed and zoomed freely.
Motion over the wire, pixels on the device
Most real-time avatars render video on a server and stream it. Spatius streams the motion data and renders in the browser. That is why the same avatar can be framed as a portrait or a close up with no extra server work, and it fits the low price per minute that Spatius publishes.
Gaussian splat avatars
The avatars are 3D Gaussian splats, not generic rigged models. An avatar ID has to point to a compatible, pre-made Spatius avatar. Handing it a photo URL is not enough.
On our own site we load a compressed copy of the model, about 30% of the full size, and share it between every frame that shows the same avatar.
Running it well
Preload the model before the call starts, unlock browser audio before the first greeting, and check that fresh motion is arriving rather than trusting that the participant has joined.
Because the track carries motion, a standard recording of the video track does not contain a visible face. Recordings need a path that captures the rendered avatar.
Trade-offs
Rendering depends on the viewer's device and browser, and the SDK has to be loaded, so it does not fit a plain video tag. There is no visual perception layer: the avatar does not see the person it talks to.