September '26 Benchmarks: Tavus vs Anam, LemonSlice and HeyGen




Tavus · Research Brief
We compared Tavus with three other real-time avatar providers, Anam, LemonSlice and HeyGen, in two ways: blind studies in which 309 people chose between clips, and seven standard industry video metrics. In every comparison, Tavus and the other provider animated the same person saying the same words.
Based on the studies, people preferred Tavus in all three matchups, and Tavus had the best average score on all seven metrics, although not every difference was statistically significant.
The findings:
We built this benchmark to be fair and standardized, with methodology established before any testing began. Every provider ran out of the box on the default settings of its own integration, and we tuned none of them, including Tavus. In every comparison, Tavus and the other provider animated the same person with the same audio, so the avatar was the only thing that differed. The audio covered five language conditions: two variants of English, plus Spanish, French and Hindi.
We measured the providers in two ways. Two blind studies asked people which avatar looked and sounded more natural, and seven standard metrics scored each avatar's face likeness, realism, lip sync and head motion. The study was not pre-registered, but the rating question and the exclusion rule were built into the study app and applied automatically, and we report every result.
Tavus can create an avatar from a short video of a person. Anam and LemonSlice create avatars only from a photo, and HeyGen accepts video only when the person records it themselves, so all three used a photo of the same person, and their models supplied the motion. Where we had only a photo, Tavus used that photo too.
Real-time avatars are built for live conversation, but a blind, randomized comparison needs both clips in a pair to show exactly the same conversation, and live conversations vary with the language model, the voice and turn-taking. So we had each provider animate the same person speaking the same audio, recorded the results and showed raters each pair of clips in random order. That means this study focuses on visual quality and doesn't measure latency or a full conversation. An earlier study compared Tavus and Anam in live conversations, where 62.5% of 80 participants preferred Phoenix-4 Pro over Anam's Cara-3. We'll soon share a separate benchmark that evaluates end-to-end, real-time conversational performance.
We recruited 309 raters through Prolific, an outside research panel, for two separate studies with no overlap: 155 raters compared Tavus with Anam and LemonSlice, and 154 compared Tavus with HeyGen. Together they judged 3,037 pairs of clips that included Tavus. Provider names were never shown, the order of the clips was randomized, and every rater saw the same question: "Which AI avatar looks more natural?"
Each rater judged 12 pairs, and each pair worked the same way.
We counted only the pairs where a rater picked a winner. For each rater, we calculated the share of those pairs that Tavus won and averaged those shares across raters, so every rater counts equally, then tested whether the average differed from 50%. The app would have excluded any rater who skipped more than three pairs, but no one did. In the first study, 4 of each rater's 12 pairs compared Anam with LemonSlice; this report covers only the pairs that include Tavus.
For the metrics, all four providers ran through Pipecat, an open-source framework for real-time video agents, each using its own integration. Each provider ran 29 matched sessions of about 55 seconds, covering six faces and five language conditions, and every session was scored the same way after skipping the first seven seconds of start-up. We compared Tavus with each provider session by session using a paired Wilcoxon test and call a difference statistically significant when p is below 0.05.
| Standard Industry Metric | What it measures | Better when |
|---|---|---|
| Face likeness (CSIM) | How closely the avatar resembles the real person | ↑ Higher |
| Video realism (FVD) | How closely the video, including motion, resembles real video | ↓ Lower |
| Image realism (FID) | How closely individual frames resemble real video frames | ↓ Lower |
| Lip-sync confidence (LSE-C) | How well the lips match the audio, according to the SyncNet model | ↑ Higher |
| Lip-sync distance (LSE-D) | How far apart lip movement and audio are, according to SyncNet | ↓ Lower |
| Head-motion rhythm (Beat Align) | Whether head movement follows the rhythm of speech | ↑ Higher |
| Lip sync by speech model (AVSR) | Whether the lips match the words, scored by an audio-visual speech model | ↑ Higher |
Every provider was tested on the real-time model it offered when the clips were generated, with default settings.
| Provider | Model | How it ran | Video format |
|---|---|---|---|
| Tavus | Phoenix-4.5 | Tavus's Pipecat integration, and Tavus's own pipeline for the first study's clips | 16:9 (1280×720) |
| Anam | Cara-4 | Anam's Pipecat integration | 3:2 (1152×768) |
| LemonSlice | LemonSlice-2.1 | LemonSlice's Pipecat integration | 2:3 (368×560) |
| HeyGen | LiveAvatar | LiveAvatar's LiveKit integration, connected to Pipecat | 16:9 (1280×720) |
During testing, we found and fixed a bug in our own Pipecat integration, and the Tavus results use the fixed version, which is now in production.
Which avatar people preferred
People preferred Tavus in all three matchups, and most individual raters did too. Counting raters rather than pairs: against Anam, 85 of 151 raters chose Tavus more often, 43 chose Anam more often and 23 were even. Against LemonSlice, the split was 128 to 14, with 12 even, and against HeyGen it was 128 to 12, with 13 even. By language, Tavus led in all five language conditions against LemonSlice and HeyGen, and in four of the five against Anam.
Standard metrics
| Metric | Better when | Tavus | Anam | LemonSlice | HeyGen |
|---|---|---|---|---|---|
| Face likenessCSIM | ↑ Higher | 0.926 | 0.886 | 0.775 | 0.902 |
| Video realismFVD | ↓ Lower | 106.5 | 155.4 | 210.4 | 160.4 |
| Image realismFID | ↓ Lower | 24.6 | 27.8 | 43.5 | 40.2 |
| Lip-sync confidenceLSE-C | ↑ Higher | 8.52 | 8.29 | 7.48 | 7.67 |
| Lip-sync distanceLSE-D | ↓ Lower | 7.84 | 7.85 | 8.76 | 8.41 |
| Head-motion rhythmBeat Align | ↑ Higher | 0.469 | 0.446 | 0.453 | 0.456 |
| Lip sync by speech modelAVSR | ↑ Higher | 0.268 | 0.250 | 0.218 | 0.147 |
Tavus had the best average on all seven metrics. FVD and FID are calculated over each provider's full set of videos (24 per provider) rather than per session, so they have no significance test. We also measured how much each avatar's head moved: Anam's avatars moved the most (0.645), followed by LemonSlice (0.608), Tavus (0.483) and HeyGen (0.398). We don't rank this, because more movement isn't better or worse on its own. What matters is whether movement fits the speech, which head-motion rhythm measures, and Tavus had the best average on that metric.
Phoenix-4.5, the model tested here, is available now in Tavus PALs: AI humans that see, hear and respond face to face in real time. Teams use PALs for:
You can build a PAL without code in PAL Maker, build one into your own product with the developer API or through the same Pipecat integration we used for this study's metrics, or have the Tavus team build and run one for you as an enterprise solution. Start building for free, read the docs or talk to our team.