Now in beta

Your voice agentalready talks.Now give it a face.

FaceMode turns your agent's audio into a live, lip-synced avatar.
No new voice stack. No video pipeline to build.

Real-time lip syncLiveKit + Daily.co compatibleAny TTS or voice agent1K resolution
Live Session
● REC
POST /api/sessions
{
"avatarId": "avatar-456",
"room": {
"type": "livekit",
"url": "wss://your-livekit.example.com",
"token": "<your room token>"
}
}
201 CREATED
{
"session": { "id": "session-123", "status": "ACTIVE" },
"ingestion": {
"ready": true,
"url": "wss://worker-01.facemode.io/ws/session-123",
"wsToken": "facemode.eyJhbGciOi..."
}
}
Streaming950 ms

[ Audio sources ]

Stream the voice you already use.

FaceMode Works with
  • ElevenLabs
  • Cartesia
  • Deepgram
  • Sarvam
  • Gnani
  • OpenAI
  • Gemini
Or

Bring any voice pipeline. FaceMode does the rest.

[ How it works ]

Audio goes in. A face comes out.

From the first audio frame to a rendered face on screen - four stages, one continuous low-latency path.

Audio in
from your TTS or agent
FaceMode
lip-synced video frames
Your room
LiveKit or Daily.co
On screen
your viewer SDK
01

Create a session

One REST call returns the room binding and a one-time credential for the audio socket.

02

Stream your audio

Pipe the audio your agent already produces, about 100ms at a time. No client-side resampling.

03

The face renders live

FaceMode generates lip-synced video frames and publishes them into your room as an ordinary participant.

04

Your app just watches

Subscribe with the LiveKit or Daily.co client SDK you already ship. Nothing new on the frontend.

[ Setup prompt ]

One prompt. Your agent does the wiring.

Paste it into the coding agent you already use. It reads the FaceMode docs, checks your project, and adds FaceMode with our LiveKit plugin, Pipecat plugin, or direct API. It shows you the plan before it edits anything.

Read the integration guides
PromptPaste into your coding agent

Add FaceMode's avatar layer to this project's existing TTS or voice agent. Keep my voice stack and conversation behavior. ...

[ What you get ]

Keep your stack. Add a face.

FaceMode does one job. It turns the audio your agent already produces into a talking avatar, in the room you already own.

Real-time lip sync

Audio in, a lip-synced face out, as one continuous stream in 1K resolution. No render jobs, no polling, nothing to wait for.

01

Fast enough to interrupt

A 960ms target from audio to frame, so people can cut in mid-sentence and the face keeps up.

02

Your room, your rules

FaceMode joins the LiveKit or Daily.co room you already own and publishes as one more participant. Your viewers subscribe with the client SDK you already ship.

03

Any voice you already use

ElevenLabs, Cartesia, Deepgram, Sarvam, Gnani, OpenAI, Gemini, or your own pipeline. If it can stream audio, it can drive a face.

04

Plugin or plain API

On LiveKit Agents or Pipecat, install the plugin and keep writing agent code. On anything else, create a session over REST and stream audio. Any language that holds a socket open works.

05

Never a frozen face

Between utterances the avatar keeps blinking and breathing, then crossfades into speech. No freeze-frame, no state handling on your side.

06

[ Why a face helps ]

People respond to a face.

Add a face where customers already look - training, intake, guidance, and live help. Every claim below comes from published research.

Healthcare, randomised controlled trial

0.0%

Knowledge gain, against 3.7% without

0.0%

Patients satisfied with the session

Patients taught by an avatar gained six times more knowledge than a control group - and nearly all of them said they were satisfied.

Journal of Advanced Nursing

And in every other sector studied

[ Avatars ]

Not every brand wants a photoreal human.

Pick a face from the library or upload your own portrait. FaceMode lip-syncs all five styles the same way.

Realism

Editorial-grade human portraits.

Semi-realism

Human, gently stylised.

3D animation

Smooth, warm, character-friendly.

Anime / manga

Cel shading and crisp linework.

Flat vector

Minimal shapes, brand-safe.

Bring your own face. Upload one front-facing portrait and it becomes an avatar.

[ Private Beta ]

Join the beta.

We are letting teams in a few at a time. Approved testers get 60 minutes of avatar video to build against the full API. No card, no subscription.

Beta access

60free minutes

Approved beta testers get 3,600 credits - 60 free minutes of streamed avatar video - to build and test with the full API.

Join the waitlist
  • REST API plus your LiveKit or Daily.co room
  • Plugins for LiveKit Agents and Pipecat
  • Avatar library plus your own uploads
  • Usage dashboard in the console
3,600 credits1 credit = 1 second of video10 minutes per session1 session at a time

[ Questions ]

Before you ask.

Do I have to change my voice stack?

No. FaceMode takes the audio your stack already produces. Your STT, LLM, TTS, and turn detection stay exactly where they are.

Who owns the room?

You do. You create a LiveKit or Daily.co room and hand FaceMode a token to join and publish. FaceMode never sits between your app and your room.

I am already on LiveKit Agents or Pipecat.

Install the plugin. It creates the session, opens the audio socket, and reconnects on drops. Python and JS for LiveKit Agents, Python for Pipecat.

I am on neither.

Create a session over REST, open a WebSocket, stream audio. Any language that can hold a socket open works.

Can people interrupt the avatar?

Yes. Stop sending audio and the face crossfades back into its idle loop instead of freezing, then picks straight back up.

Can I use my own face?

Upload one front-facing portrait and it becomes an avatar. Or pick from the library across five art styles.

How does billing work?

Credits, not seats or subscriptions. One credit is one second of streamed avatar video. FaceMode is in beta right now - top-ups for extra credits open once we hit general availability.

[ Ready when you are ]

Give it a face.

Your agent already knows what to say. Spin up a session and let people watch it say it.

Any TTS SupportedWorks with Livekit-Agents & Pipecat1K resolution