Introducing Griffin
Griffin is the world’s first Human Interaction Model. It sees you, hears you and talks back, all at the same time. Not a chain of systems taking turns, but one model in one conversation: it nods along, cuts in, gives you room to think. After a one-minute video call with it, not knowing what it was, 48% of people believed they’d been talking to a real person.
.avif)
Today we’re introducing Griffin, our first Human Interaction Model (HIM). HIMs are a new class of models designed to understand and generate face-to-face, real-time human interaction. They listen while they talk, and they pay attention to expressions and pauses, not just words.
Griffin builds on our earlier research into real-time perception, conversation and expressive, human-like video. It combines those model capabilities into one unified video-to-video system.
Griffin is the first model to pass the real-time video Turing test. In a live study, 48% of participants who talked with Griffin thought they had talked with a real human. Previous systems, including Tavus’s own industry-leading Conversational Video Interface (powered by Phoenix-4.5, Sparrow-2 and Raven-1), had a pass rate of 2.4% or lower. This represents a huge breakthrough in human-machine understanding, and is only possible because of Griffin’s video-to-video duplex approach, with audiovisual generation, movement and conversational modeling all in one system, operating in real time.
Griffin-Lite, a research preview, is available today to a select group of early testers, with a wider release of a more powerful model to follow. This post covers what Griffin can do, how we built it, how we measured it, our approach to safety and what comes next.
Human computing
A message from Hassaan
Humans are evolutionarily designed to communicate face to face. We speak as much through our words as we do our expressions, tone, gestures and timing. So much of human communication is non-verbal. It’s an art, a dance.
Machines don’t understand this art. They require us to meet them where they are, to learn their language: the right command, the right button click, the right prompt.
At Tavus, we believe in a future where machines meet us where we are. Where machines learn to communicate in the way that we most effortlessly do. They learn to see, hear, respond and even look like we do. In effect, we want computing to become invisible. We want it to feel second nature.
In a great conversation, you don’t spend your time evaluating the mechanics. You think about what you want to say, what you’re hearing and seeing, and ultimately the conversation just flows. In great conversations, you aren’t thinking about the conversation at all.
With machines, those mechanics are distracting and demand your attention.
Did it hear me? Did it understand me? You share something personal or difficult, and the face continues to smile, or the faceless, cold machine sits there, then responds with a monotonous “I’m sorry to hear that.” You pause to collect your thoughts, and it jumps in and responds to incomplete context. It needs time to answer, but nothing in its expression or voice or gesture tells you it’s working on a response.
Each of these moments is you managing the machine. You adjust your timing, simplify your language, make sure to think complete thoughts before saying anything. You remind yourself: don’t leave dead air! Remember to be very clear about what you want! You lower your expectations. The system, in all its intelligence, might be capable of infinite incredible work. It can discover new mathematics, dammit! Yet getting help for the simple stuff takes more effort than you have to give.
Human communication is an art, a dance. The right move, at the right time. A deep understanding of what you mean, when the words themselves could mean many things. A nod of acknowledgment to signal understanding. An expression that tells you everything without saying anything.
All of these are things we do as humans without thinking about it. These signals let your attention stay on the conversation, not how to converse. As a machine becomes more coherent at expressing and understanding these signals, communicating with it, too, becomes something you don’t have to think about.
Griffin watches and listens while it talks.
Griffin-Lite is a preview of our first Human Interaction Model: a full-duplex video-to-video model that responds to human behavior in real time. Perception, deciding when and how to respond, and expressive speech and video generation all happen at the same time.
- 1It says “mm-hm” while you’re still talking.
- 2It answers without leaving dead air.
- 3It stops the moment you cut in.
- 4It reacts to something it sees.
The timing in this diagram is illustrative and was not measured from a real session.
It holds up its end of the conversation.
To see what it looks like in practice, we brought people who had never used Griffin together with Tavus researchers for open-ended conversations and short demonstrations.
You hear the reaction and see it in the same instant.
Griffin has learned to express behaviors, emotions and gestures according to conversational context. It laughs, changes its tone and shifts in response to what the other person says and does.
Rudy tells Vanessa he’s been promoted. She reacts before he’s finished, and the two of them talk over each other for a moment without losing the thread.
The whole frame is generated, not just the face.
Griffin generates every pixel of every frame in real time from one reference image, which lets it control the entire body with emergent gestures, including arms, hands and fingers.
A round of Simon says with Mars. Griffin copies the gesture only when the person says “Simon says,” and calls their bluff when they don’t. Every movement is generated on the spot, body and background included.
You don’t have to wait for your turn.
Griffin continuously evaluates the state of the conversation, independently of conversational turns. It can interrupt, adjust, backchannel or be interrupted without losing its place.
Ari and Griffin make up a story together, cutting each other off as they go, and Griffin runs with every twist.
It notices what’s in front of the camera, even while it’s talking.
Griffin takes dialogue beyond verbal understanding, embedding and reacting to visual context and awareness when appropriate.
Griffin coaches Ari through a Rubik’s Cube. It watches the cube as he turns it, keeps acknowledging him while he works, and when he pauses mid-solve it waits until he’s ready.
It knows how long it’s been.
Griffin understands time as part of the conversation: how long a silence has lasted, what a sustained silence means, and when to speak again on its own.
Sagar solders a motherboard while Griffin guides him. It keeps track of time and where he is in the job, and it speaks up when the next step is due rather than whenever he goes quiet.
Griffin learns the whole conversation instead of one part of it.
Perception, conversational decision-making, and generation are in continuous interplay throughout human-to-human interaction, and we designed Griffin around the same principle. Central to our approach is a conversational model that controls speech and nonverbal behavior directly, with streaming speech and video generators that follow its decisions as they change.
Griffin is a two-part system. A Continuous Conversational Modeling engine perceives the incoming audio and video, decides when and how to respond, and produces signals that determine what should be said and how. An Audio-Visual Generation engine, made up of a Streaming Speech Generation and a Streaming Video Generation architecture, converts those signals into speech and video. Perception, decision-making, and generation run concurrently throughout the conversation, so the model can respond to changes while it is speaking as well as while it is listening.
Griffin makes conversational decisions at regular sub-second intervals rather than once per turn. At each interval it assesses the state of the conversation, including what has been said and the user’s verbal and nonverbal behavior, and decides what to say next and how to say it. This differs from cascade systems, which chain speech recognition, a language model, and speech synthesis one after the other, and which wait for the user to finish speaking before they begin to respond.
The output controls not only what is said but also how the speech is delivered and how the model behaves nonverbally, including emotional tone, stance, facial expression, and gesture. Because these decisions are made continuously and are conditioned on context, the model can take and yield the turn on the basis of what is being said rather than when the audio stops, so a pause for thought is not treated as the end of the turn. The same mechanism lets the model nod in agreement while the user is speaking, backchannel and confirm at natural points, and adjust the emotional content of its response to the user’s speech and nonverbal behavior.
Additionally, Griffin perceives video as well as audio, where most interactive systems perceive audio alone. Visual input lets the model read the user’s nonverbal communicative signals, such as gaze and facial expression, and it also gives the model access to the user’s environment and to visual material the user chooses to share, such as their screen. The emitted outputs are converted to streaming audio and, together with streaming control signals, drive video generation, as explained below.
Fast and expressive streaming speech generation
Griffin-Lite’s speech generation produces high-quality speech at conversational latency and can clone a speaker’s voice from about 10 seconds of audio. The model is a fast autoregressive diffusion transformer (VDiT). It takes an encoding of the reference clip as a prefix and generates the new speech progressively, one latent chunk at a time, as increments of outputs and controls arrive from the conversational model, instead of waiting for the full utterance.
A compact continuous codec
Griffin’s speech generation speed and fidelity start with how it represents sound. It relies on Tavec, a convolutional autoencoder that maps 48 kHz audio into a compact continuous latent: 40 values per frame, 100 frames per second, and no codebooks. Its fully causal (streaming-ready) decoder carries state across chunks, turning streamed latents into one seamless waveform, with audio packets as small as 10 ms.
This continuous representation supports straightforward regression-based flow matching and fully differentiable audio pipelines, without quantizer gradient approximations. Its compactness also makes long sequences manageable: a minute of speech corresponds to 6,000 latent frames, against 2.88 million samples of raw audio.
Fast and expressive streaming video generation
The other side of Griffin’s audio-visual generation system is a fast diffusion-based video generation architecture. Previous Tavus models relied on extensive 3D priors, which limited what they could express: complex, human-like behavior such as hand and large body gestures was out of reach, and so were dynamic backgrounds.
We built it to meet three requirements: respond quickly to each incoming chunk of audio and expressive controls, handle controls that change fast, and hold visual quality, lip sync, and identity consistency over long generations. Autoregressive latent video models typically compromise on at least one of these. The architecture we arrived at generates 720p video in 320 ms chunks in real time, one latent at a time, and accepts streaming controls that let the conversational model direct nonverbal behavior such as gesturing, looking away, or changing emotion.
To build it, we distilled a large, bidirectional, many-step diffusion model into a few-step autoregressive generator. The generator takes a reference image together with streaming audio and controls, and produces one latent at a time. The distillation ran in three stages. First, we distilled the teacher into a few-step student with Distribution Matching Distillation. This student is fast but not autoregressive; it generates a long video as a single chunk. Second, we converted it into an autoregressive model with teacher forcing, so as to arrive at an architecture that can generate one latent at a time in a few steps. Third, we trained with Self-Forcing so that the model can run long autoregressive rollouts without drift, and we added a mechanism acting on the history frames, the previously generated frames the model continues from, to improve stability.
Nearly half couldn’t tell it was AI.
Ahead of this preview, we ran Griffin-Lite through a set of evaluations at different levels, from a single component to a live conversation: the video generator on its own, against published streaming diffusion models, on latency, visual quality and lip sync; the full system on VideoFDB, NVIDIA’s leading industry benchmark for full-duplex audio-visual conversation, which scores how a model reads a person’s behavior and how it responds; and the full system in live one-minute video calls with people who did not know they were talking to a model.
Video generation
We compared the video generator of Griffin-Lite against four published streaming diffusion models in an audio-to-video setting, where each model is given speech and a reference image and has to produce the talking face. Griffin-Lite produces one latent at a time and does not wait for future audio, so there is little delay between a piece of audio arriving and its effect appearing on screen. On H100s this averaged 0.43 seconds, half that of the next fastest method. In a conversation, this significantly cuts down the delay before the face reacts when the person interrupts or when a nod is due.
Griffin-Lite also scored highest on visual quality. It led every baseline on DOVER and FID, the two standard measures of video quality, and on THEval, a recently published framework built for talking heads. On lip sync it placed second on LSE-C, at 7.27. LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.
Full-duplex conversation
NVIDIA’s VideoFDB is a research and industry benchmark for evaluating the naturalness and accuracy of full-duplex audio-visual conversations. It examines how a model responds to and produces the signals of human interaction, including dialogue, gaze, facial expression, and body movement, measuring both conversation quality and response time. The benchmark scores a model on two tracks, generation and perception, each with its own rubric and its own leaderboard. NVIDIA conducts the evaluation independently using the published metrics and its own judge. Griffin-Lite scored highest on both tracks against the benchmark’s published baselines, which include commercial and open-source systems, and it is the only case where the same model is evaluated on both tracks.
The generation track tests whether the model produces the right behavior. Given a person’s audio, it scores the model’s speech and video together for fluency, affect matching, and whether the nonverbal behavior fits the moment, laughing with the person rather than after them. Griffin-Lite scored 3.83, over 1 full point ahead of the next-highest system and only 0.09 below the human reference at 3.92, which is more than 12x closer to human performance than any other system evaluated.
VideoFDB’s perception track tests the other side of the conversation, which is whether the model understands the moment. Given a person’s audio and video, it scores the model’s spoken response for fluency, conversational flow, and visual grounding, meaning whether the model uses what it sees, a pause with a glance away, for instance, rather than reacting to the words alone. Griffin-Lite scored 3.73, 0.29 points ahead of the strongest second best and 0.47 below the human reference at 4.20, marking Griffin-Lite as the most perceptive model on the benchmark among the 15 models evaluated, ahead of every frontier realtime model on the leaderboard.
VideoFDB also reports takeover-rate alignment, which measures how closely the model’s decisions about when to speak match the timing in the reference conversations. Griffin-Lite scored 62.8% on generation and 73.8% on perception, the highest of any system on both tracks.
Face-to-face study
Quantitative benchmarks are important, but some aspects of naturalness are difficult to measure. For every new model, we run a study with participants recruited through an independent research platform to understand how human-like it is in a real, live conversation. Participants have a short video call with the model without being told what it is, and afterward we ask whether they think their partner was a real person. Until now, almost no one has said yes.
On Phoenix-4.5, 2.4% of participants (n=41) believed their partner was a real person. On Griffin-Lite, nearly half did: 48% (n=54). This represents a milestone: to the best of our knowledge, the first model to have ever passed the video Turing test.
Methodology
Participants in the US and Europe were told they would be matched with another participant for a one-minute video call to discuss what they were looking forward to this year. Their partner was in fact a PAL powered by Griffin-Lite, generating her face, voice, and responses in real time. After the call, participants were asked to write down their partner’s answer to the question and to rate their partner on several axes around naturalness and trustworthiness. They were also asked to rate the conversation itself: how well it flowed, and whether they felt their partner was really listening. Only at the end of the survey were participants asked whether it had crossed their mind that their partner might not be a real person, and if so, when. Every participant was then told that their partner had been an AI model.
Results
Of the 54 participants who spoke with Griffin-Lite, 26 believed their partner was a real person. Participants on both sides were confident in their answer: those who said real averaged 79% confidence, and those who said AI averaged 81%. Over half of participants said the possibility had not crossed their mind during the call, and nearly all of that group went on to say their partner was real. Those who did suspect tended to suspect within the first 20 seconds.
Griffin-Lite was rated favorably on every axis. On a 7-point scale, participants on average rated it a 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether they would enjoy talking with it again, a score that held at 5.4 among participants who said it was AI. The conversation scored 5.5 for whether participants felt their partner was really listening and 4.9 for flowing naturally, the lowest of the five.
A model people can mistake for a person has to be released carefully.
Given Griffin-Lite is the first model to pass the Turing test, we understand the responsibility to be thoughtful on the risks and release strategies.
The same properties that make Human Interaction Models powerful interfaces for natural communication between human and machine allow them to deceive a human into believing it is not AI.
We believe further alignment and safety procedures are required for safe release, and we are working on safe disclosure features, as well as with organizations tackling AI safety. Contact us if you’d like to participate in these evaluations and testing. Griffin-Lite will not be available for use for customers at this time, though it is available for select trusted testers as a research preview. We anticipate releasing Griffin very soon after these safety concerns are addressed.
This is the start of a new kind of model.
Griffin is an early step toward computers that people can work with instead of operate. As these models improve, they can bring knowledge and emotional understanding together and play a larger part in how people learn, work, and get help.
For a student, that could mean a tutor who notices when an explanation isn’t making sense and tries another one. For someone at work, it could mean rehearsing a hard conversation with a counterpart who reacts the way a real person might. For a customer, it could mean holding a broken part up to the camera and working out the fix together, without knowing what the part is called.
In each of those cases, the person doesn’t have to turn what they need into the right command or menu first. They can explain it the way they would to another person, and the model can work out what they mean from their words, their face, and the moment. That’s what we mean by human computing, and Griffin is our biggest step yet toward it.
Credits
Griffin was built by the Tavus research and engineering teams [names to come]. Our thanks to Baseten, Daily and Cerebrium for the infrastructure behind the preview, to Queen Mary University of London for research collaboration, and to NVIDIA for building and scoring the Video Full-Duplex Benchmark.
Build with Tavus today
Griffin isn’t on the Tavus platform yet. It’ll come once we’ve worked out how to release it safely. Until then, you can build PALs, AI you talk to face to face, on the models 150,000 developers and businesses already use: Phoenix generates the face, Raven sees and understands you, and Sparrow knows when to talk.