Tavus · Research Brief

We compared Tavus with three other real-time avatar providers, Anam, LemonSlice and HeyGen, in two ways: blind studies in which 309 people chose between clips, and seven standard industry video metrics. In every comparison, Tavus and the other provider animated the same person saying the same words.

Based on the studies, people preferred Tavus in all three matchups, and Tavus had the best average score on all seven metrics, although not every difference was statistically significant.

TestTwo blind studies in which people compared clips side by side, plus seven standard video metrics.
Sample309 people judged 3,037 pairs of clips, and each provider ran 29 matched sessions for the metrics.
ModelsWe tested Tavus Phoenix-4.5, Anam Cara-4, LemonSlice-2.1 and HeyGen LiveAvatar.
DisclosureTavus designed and ran the study. We disclose our methodology and report every result, including those that don't favor Tavus.

Key statistics

Preferred Tavus over Anam
60.8%
95% CI 55.3% to 66.2%
p < 0.001
Preferred Tavus over LemonSlice
85.1%
95% CI 81.1% to 89.1%
p < 0.001
Preferred Tavus over HeyGen
75.0%
95% CI 71.8% to 78.2%
p < 0.001
Metrics where Tavus scored best
7 of 7
8 of 15 differences significant.
All 8 in Tavus's favor

The findings:

  • People preferred Tavus over all three providers. On average, raters chose Tavus 60.8% of the time against Anam, 85.1% against LemonSlice and 75.0% against HeyGen.
  • Tavus had the highest face likeness of the four providers, and the difference was statistically significant against each of them. Tavus also had the best video- and image-realism scores.
  • Tavus's lip sync was significantly better than LemonSlice's on all three lip-sync measures and HeyGen's on two of the three.
  • Tavus's lead over Anam came from faces where Tavus could use a short video, which the other providers could not.

Methodology

How we tested

We built this benchmark to be fair and standardized, with methodology established before any testing began. Every provider ran out of the box on the default settings of its own integration, and we tuned none of them, including Tavus. In every comparison, Tavus and the other provider animated the same person with the same audio, so the avatar was the only thing that differed. The audio covered five language conditions: two variants of English, plus Spanish, French and Hindi.

We measured the providers in two ways. Two blind studies asked people which avatar looked and sounded more natural, and seven standard metrics scored each avatar's face likeness, realism, lip sync and head motion. The study was not pre-registered, but the rating question and the exclusion rule were built into the study app and applied automatically, and we report every result.

Tavus can create an avatar from a short video of a person. Anam and LemonSlice create avatars only from a photo, and HeyGen accepts video only when the person records it themselves, so all three used a photo of the same person, and their models supplied the motion. Where we had only a photo, Tavus used that photo too.

Real-time avatars are built for live conversation, but a blind, randomized comparison needs both clips in a pair to show exactly the same conversation, and live conversations vary with the language model, the voice and turn-taking. So we had each provider animate the same person speaking the same audio, recorded the results and showed raters each pair of clips in random order. That means this study focuses on visual quality and doesn't measure latency or a full conversation. An earlier study compared Tavus and Anam in live conversations, where 62.5% of 80 participants preferred Phoenix-4 Pro over Anam's Cara-3. We'll soon share a separate benchmark that evaluates end-to-end, real-time conversational performance.

The human studies

We recruited 309 raters through Prolific, an outside research panel, for two separate studies with no overlap: 155 raters compared Tavus with Anam and LemonSlice, and 154 compared Tavus with HeyGen. Together they judged 3,037 pairs of clips that included Tavus. Provider names were never shown, the order of the clips was randomized, and every rater saw the same question: "Which AI avatar looks more natural?"

Each rater judged 12 pairs, and each pair worked the same way.

Blind human studies
Every rater judged 12 pairs the same way
Each rater answered the same question for every pair: which AI avatar looks more natural?
1Clip A, with sound
The rater watches and hears the first clip.
13 seconds
2Clip B, with sound
The rater watches the same person saying the same words, animated by the other provider.
13 seconds
3Both clips, muted
The two clips play side by side without sound, and the rater picks A, B or a tie.
Provider names hiddenClip order randomizedRaters from ProlificTies not counted

We counted only the pairs where a rater picked a winner. For each rater, we calculated the share of those pairs that Tavus won and averaged those shares across raters, so every rater counts equally, then tested whether the average differed from 50%. The app would have excluded any rater who skipped more than three pairs, but no one did. In the first study, 4 of each rater's 12 pairs compared Anam with LemonSlice; this report covers only the pairs that include Tavus.

The metrics

For the metrics, all four providers ran through Pipecat, an open-source framework for real-time video agents, each using its own integration. Each provider ran 29 matched sessions of about 55 seconds, covering six faces and five language conditions, and every session was scored the same way after skipping the first seven seconds of start-up. We compared Tavus with each provider session by session using a paired Wilcoxon test and call a difference statistically significant when p is below 0.05.

Standard Industry MetricWhat it measuresBetter when
Face likeness (CSIM)How closely the avatar resembles the real person↑ Higher
Video realism (FVD)How closely the video, including motion, resembles real video↓ Lower
Image realism (FID)How closely individual frames resemble real video frames↓ Lower
Lip-sync confidence (LSE-C)How well the lips match the audio, according to the SyncNet model↑ Higher
Lip-sync distance (LSE-D)How far apart lip movement and audio are, according to SyncNet↓ Lower
Head-motion rhythm (Beat Align)Whether head movement follows the rhythm of speech↑ Higher
Lip sync by speech model (AVSR)Whether the lips match the words, scored by an audio-visual speech model↑ Higher

Models and configurations

Every provider was tested on the real-time model it offered when the clips were generated, with default settings.

ProviderModelHow it ranVideo format
TavusPhoenix-4.5Tavus's Pipecat integration, and Tavus's own pipeline for the first study's clips16:9 (1280×720)
AnamCara-4Anam's Pipecat integration3:2 (1152×768)
LemonSliceLemonSlice-2.1LemonSlice's Pipecat integration2:3 (368×560)
HeyGenLiveAvatarLiveAvatar's LiveKit integration, connected to Pipecat16:9 (1280×720)

During testing, we found and fixed a bug in our own Pipecat integration, and the Tavus results use the fixed version, which is now in production.

Results

Which avatar people preferred

Blind human studies
People preferred Tavus in all three matchups
Tavus’s average share of each rater’s decided pairs. 50% would mean no preference.
TavusThe other provider95% confidence interval
vs Anam
151 raters
60.8%
95% CI 55.3–66.2%
vs LemonSlice
154 raters
85.1%
95% CI 81.1–89.1%
vs HeyGen
153 raters
75.0%
95% CI 71.8–78.2%
Brackets show 95% confidence intervals. All three results differ from 50% at p < 0.001. Rater counts include everyone with at least one decided pair in that matchup.

People preferred Tavus in all three matchups, and most individual raters did too. Counting raters rather than pairs: against Anam, 85 of 151 raters chose Tavus more often, 43 chose Anam more often and 23 were even. Against LemonSlice, the split was 128 to 14, with 12 even, and against HeyGen it was 128 to 12, with 13 even. By language, Tavus led in all five language conditions against LemonSlice and HeyGen, and in four of the five against Anam.

Blind human studies
Most individual raters chose Tavus more often
One square per rater, colored by the provider that rater chose more often in the matchup.
vs Anam
151 raters
852343
vs LemonSlice
154 raters
1281214
vs HeyGen
153 raters
1281312
Chose Tavus more oftenEvenChose the other provider more often
Raters with no decided pairs in a matchup are not shown.

Standard metrics

Standard video metrics
Tavus had the best average on all seven metrics
Average over 29 matched sessions per provider. FVD and FID use each provider’s full set of 24 videos.
MetricBetter whenTavusAnamLemonSliceHeyGen
Face likenessCSIM↑ Higher0.9260.8860.7750.902
Video realismFVD↓ Lower106.5155.4210.4160.4
Image realismFID↓ Lower24.627.843.540.2
Lip-sync confidenceLSE-C↑ Higher8.528.297.487.67
Lip-sync distanceLSE-D↓ Lower7.847.858.768.41
Head-motion rhythmBeat Align↑ Higher0.4690.4460.4530.456
Lip sync by speech modelAVSR↑ Higher0.2680.2500.2180.147
Tavus’s column holds the best value in every row. Not every difference is statistically significant.

Tavus had the best average on all seven metrics. FVD and FID are calculated over each provider's full set of videos (24 per provider) rather than per session, so they have no significance test. We also measured how much each avatar's head moved: Anam's avatars moved the most (0.645), followed by LemonSlice (0.608), Tavus (0.483) and HeyGen (0.398). We don't rank this, because more movement isn't better or worse on its own. What matters is whether movement fits the speech, which head-motion rhythm measures, and Tavus had the best average on that metric.

Building with Tavus today

Phoenix-4.5, the model tested here, is available now in Tavus PALs: AI humans that see, hear and respond face to face in real time. Teams use PALs for:

  • Sales. AI SDRs that qualify inbound leads, run demos, handle objections and book meetings.
  • Recruiting. AI interviewers that screen candidates face to face at scale, adapt their questions to each answer and evaluate every candidate the same way.
  • Healthcare. HIPAA-compliant agents for patient intake, support and education.
  • Learning and development. Realistic simulations for training, coaching and practice, from onboarding to sales role-play.
  • Customer support. High-touch, face-to-face support that scales without adding headcount.

You can build a PAL without code in PAL Maker, build one into your own product with the developer API or through the same Pipecat integration we used for this study's metrics, or have the Tavus team build and run one for you as an enterprise solution. Start building for free, read the docs or talk to our team.