Phoenix-4.5 is here: The fastest and most expressive real-time human rendering model on the market. Learn more.
Products
build with charlie
PAL Maker
No-code way to build and deploy a PAL
explore cvi
Developer API
Build and deploy PALs with the API 
Solutions
Solutions
By Use Case
Sales Agents
Healthcare Agents
Interview Agents
L&D Agents
Custom Agents
API PLatform
API PLatform
Get Started
Developer API
Docs
API Reference
Pricing
Features
Overview
Magic Canvas
Presentation
Components
Internet Search
‍Google Meets
Memory
Knowledge Base
Enterprise
Enterprise
Book Demo
Research
Research
Explore
Research Overview
Blog
Models
Phoenix-4.5
Raven-1
Sparrow-2
pricing
LoginGet Started
get started
LoginGet Started
Select an Account Type

Choose how you want to experience Tavus. Whether you’re building with our APIs or meeting a PAL, you can switch anytime.

Developer Account

Build real-time, human-like AI experiences using Tavus APIs and tools.

Best for developers, founders, and teams integrating Tavus into a product.

NEW DEVELOPER ACCOUNT
PALs Account

Meet your personal AI companions who listen, remember, and are always present.

Best for individuals looking to talk, explore, and connect with a friend.

NEW PALS ACCOUNT

our research

A new kind of
research lab

Bridging the human-machine divide

our approach

Human communication
is like a dance

Human conversation is a rhythm—every glance, pause, and tone changes the meaning. At Tavus, we study that rhythm, designing AI that understands emotion, intent, and timing as one signal. We’re building systems that don’t just respond, they move with you.

SEE DOCS

The Dance

PIONEERING HUMAN COMPUTING

We’re teaching machines the art of being human: bringing together rendering, perception, emotion, and understanding to create AI humans that
feel natural, intuitive, and alive.

Research Directive

We’re building AI that feels human—machines that see, listen, and respond naturally.

models

Models

We build models to teach machines to see, hear, understand and even look human. We give machines presence and EQ, allowing them to understand you deeply, and in turn, building trust and connection with them.
Rendering

Phoenix [4.5]

Phoenix-4.5, the most natural and realistic human rendering model on the market, designed to give AI a face, presence, and emotional range. The A gaussian-diffusion based model synthesizes high-fidelity facial behavior in real time, with contextually accurate emotions, expressions, and movement.
The Eyes and Ears

Raven [1]

Raven-1, our multi-modal perception model, giving machines the ability to see, hear and understand in real-time. It translates facial expressions, tone, gaze, emotion, and environmental context into rich conversational signals, helping your AI respond with empathy, awareness, and intent.
The Rhythm

Sparrow [2]

Sparrow-2 is our real-time conversational understanding model, built for natural conversational flow in noisy, multi-speaker environments. It models semantics, prosody, speaker identity, backchannels, interruptions, background speech, noise, and unclear audio to determine when to listen, wait, speak, or continue speaking.

Research areas

We study how intelligence perceives context, emotion, and tone to create AI that understands and acts as humans do.

contextual perception

Understanding meaning beyond words. Tone, timing, intent, and everything unsaid.

Audio understanding

Teaching machines to truly listen. Not just to sounds, but to emotion, cadence, and rhythm.

Agentic interaction

Building systems that act with awareness, not automation. Capable of response, reasoning, and restraint.

human-like speech

Synthesizing voice that carries emotion, not just words. Warmth, hesitation, humor, humanity.

Real-time rendering

Turning intelligence into motion. Seamless, lifelike expression that feels natural and alive.

Conversational intelligence

Making dialogue intuitive and human. Conversations that adapt, remember, and build trust over time.

CVI Terminal

Read our latest research

We study how intelligence perceives context, emotion, and tone to create AI that understands and acts as humans do.

view all

Research

Phoenix-4.5: A New State of the Art in Real-Time Human Rendering

Phoenix-4.5 is the fastest and most expressive real-time human rendering model on the market, and the closest AI has ever come to passing the Turing test face to face. The whole PAL moves with every word, in 134 ms from audio to video.

Minh Anh Nguyễn

9.10.2026

view all

Research

Sparrow-2: Beyond Turn-Taking to Whole-Scene Conversational Understanding

Sparrow-2 is a real-time conversational understanding model that jointly models turn-taking, interruptions, backchannels, and the full acoustic scene as a single, unified system. It is an audio-native, streaming-first engine, rebuilt from the ground up, that goes beyond endpoint detection to transform the entire audio stream into decisions about when to listen, wait, or speak at a native 10 ms frame rate.

Brian Johnson

8.27.2026

view all

Research

40 Years Later: Fulfilling Knowledge Navigator's Promises

Forty years ago, Apple imagined the Knowledge Navigator. Meet Dom, our real-life take on it, and the human computing interface from Tavus that powers him.

Hassaan Raza

6.16.2026

view all

Research

Phoenix-4: Real-Time Human Rendering with Emotional Intelligence

Phoenix-4 is the first real-time model to generate and control emotional states, active listening behavior, and continuous facial motion as a single, unified system. It is a real-time behavior generation engine, built from the ground up, that goes beyond photorealism to transform conversation data into emotionally responsive, context-aware facial expression and head motion with millisecond-level latency.

Eloi Du Bois

2.18.2026

view all

Research

Raven-1: Bringing Emotional Intelligence to Artificial Intelligence

Introducing Raven-1. A multimodal perception system that captures not just what users say, but how they say it, how they look when they say it, and what that combination actually means. It interprets tone, expression, hesitation, and context in real time, enabling AI that can truly understand intent rather than simply respond to words.

Mert Gerdan

2.10.2026

view all

Research

Sparrow-1: Human-Level Conversational Timing in Real-Time Voice

Sparrow-1 is a specialized, multilingual audio model for real-time conversational flow and floor transfer. It predicts when a system should listen, wait, or speak, enabling response timing that mirrors human conversation rather than simply responding as fast as possible.

Brian Johnson

1.13.2026

view all

Research

Introducing the evolution of Conversational Video Interface – now with Emotional Intelligence

Introducing our new family of state-of-the-art AI models: Phoenix-3, Raven-0, and Sparrow-0. Together they bring Conversational Video Interfaces (CVI) to the next level, and power Charlie, our new demo persona.

Julia Szatar

3.6.2025

view all

Research

Phoenix-2: Advanced Techniques in Talking Head Generation — 3D Gaussian Splatting

This paper will cover the past, present and future of the talking-head generation research field. Specifically, we will dive deep into the trending 3D scene representations (NeRF -> 3DGS) and the benefits of employing 3DGS in avatar applications.

Christian Safka

7.24.2024

view all

Research

Sparrow-0: Advancing Conversational Responsiveness in Video Agents with Transformer-Based Turn-Taking

In this paper, we dive into the development and research behind Sparrow-0, exploring the innovative transformer-based approach for turn-taking and its integration alongside Raven and Phoenix models within our Conversational Video Interface (CVI), an end-to-end operating system designed for building responsive video agents.

Brian Johnson

4.2.2025

view all

Research

Phoenix-1: Realistic Avatar Generation in the Wild

This research paper, written by the Tavus team, details the development of Phoenix, a groundbreaking generative model for realistic avatar creation and text-to-video generation. Phoenix leverages audio and text-driven 3D models, integrating volumetric rendering techniques and 2D Generative Adversarial Networks (GANs) to create lifelike replicas from short video clips.

Christian Safka

4.1.2024

See all research

Ethical and aligned
by design

We believe technology earns trust through honesty, not opacity. Tavus is built on informed consent, transparent systems, and full disclosure—no fine print, no hidden levers. Every model, dataset, and likeness we use exists with permission and purpose. You deserve to know how the magic works, and we’re here to show you.

Learn more

Where research becomes reality

Our research manifests as the traits that make AI feel human.

EXPRESSIVE
Empathetic
Actionable
Personalized

Expressive
(and authentic)

AI Humans bring face-to-face connection to every conversation.

Get Started for free

benefit [1]

Real-time conversation

Trained on millions of conversations to deliver smooth, humanlike dialogue.

benefit [2]

Superhuman perception

Understands actions, emotions, and screenshares to respond with context.

benefit [3]

Lifelike
presence

Displays expressive reactions and movement that build trust and engagement.

Perceptive (and aware)

AI Humans are modeled after us: they see, sense, and understand to build trust through real conversation.

Get Started for Free

benefit [1]

Perception

Deciphers nonverbal cues like body language and micro-expressions. Uses context to adapt responses and create meaningful, two-way interactions.

benefit [2]

Multimodal

Every input adds context, ensuring the AI Human sees the full picture: screenshare, voice, and surroundings.

benefit [3]

Awareness

Monitors key events and behaviors to trigger function calls while continuously sensing subtle background shifts with real-time data.

Thinking (with agency)

AI Humans are fully formed, with the cognitive skills needed for efficient, effective conversations.

Get Started for Free

benefit [1]

Knowledge

Industry-leading RAG grounds responses in your data. 15x faster than other solutions.

benefit [2]

Memory

Remembers past interactions to personalize responses and pickup conversations where they left off. Free to toggle on or off to fit any interaction.

benefit [3]

Structure

Uses customizable frameworks and logic branching to naturally structure conversations and keep moving toward your goals. 

Deployable (and customizable)

AI Humans are designed to work for you: scalable, flexible, and ready to perform.

Get Started for free

benefit [1]

Scale 

Deploy and manage AI Humans at scale, with infrastructure, WebRTC, VAD, and ASR fully managed behind the scenes.

benefit [2]

Insights 

Transcripts, visual context, and emotional markers from every conversation are accessible and used to inform improve user experiences.

benefit [3]

White-labeled

Developer first APIs. With simple, plug-and-play endpoints, you can embed AI Humans into any website or platform with ease.

Join the team
decoding conversation

Join the team shaping how humans and machines understand each other. We’re researchers, engineers, and artists building AI that listens, learns, and connects like people do. If you care about the future of intelligence and how it feels, you’ll fit right in.

see Careers
PAL Maker

Bring PALs into your next AI conversation.

Get started
developer api

Bring human connection to every AI interaction.

Explore cvi

company

PricingEnterpriseCareersPartnerships

Resources

BlogPerspectivesBrand kit (download)Press kitInfo for AIs

developers

DocsAPI referencePAL MakerQuickstartllms.txt

research

Turn TakingRenderingLLM ThinkingSee all research

socials

LinkedInXDiscord

legal

ADAPrivacy policyTerms of serviceWebsite terms of serviceCharlie's PALs supplemental termsAcceptable use policy

Support

DiscordEmail support@tavus.ioSupport centerTrust center
explore with ai:
© 2026 Tavus   |  THE HUMAN COMPUTING COMPANY  |  All Rights Reserved