How to create an AI clone from 2 minutes of video




The people an organization most wants in a conversation often have the least room on their calendar. A founder closes every deal she joins. A clinician earns trust in a single visit. Their value shows up in conversation, and each conversation takes time.
An AI clone learns how a specific person looks and sounds from about two minutes of video, then carries that presence into more conversations than a calendar could ever hold.
It can hold a live exchange or read a script aloud. The result depends on decisions made before the camera turns on, on which category of system receives the footage, and on the consent that separates a legitimate clone from impersonation.
An AI clone is a software model of one specific person, built from a short clip of video and audio. It reproduces a person's face, voice, and behavioral patterns, and once trained, it can say things its subject never recorded.
Consent separates a legitimate clone from impersonation. EU Article 50 has required disclosure of AI-generated or manipulated likenesses since 2 August 2026.
Tavus, the human computing company, builds Personified Application Layers (PALs) for live exchange. A PAL is a real-time application you talk to and build a relationship with, one that sees, hears, remembers, and responds face-to-face across text, voice, and video, rather than a chatbot you ping for one-off answers. That live presence starts with what the model can learn from the training footage.
A clone can do two things with a script: play it back as a finished video, or use it as raw material for a live conversation. Playback works for an announcement or a training module. The live option is what most enterprise teams want, because it produces outcomes a static clip cannot.
A PAL delivers that live presence through a closed loop of models working together, and the practical wins show up in the conversation itself:
Those capabilities are what a static clip cannot deliver, and what makes a two-minute recording worth training in the first place.
Two minutes is enough for current Replica training paths to capture what makes one person recognizable on video. The clip pairs a speaking segment with a still segment, and each half teaches the model something specific.
From that recording, the model learns:
Capturing all of this cleanly depends on how you set up the recording and what the subject does on camera.
In PAL Maker, the no-code setup flow walks you through uploading the clip, selecting a training path, and previewing the resulting Replica besfore it goes live. The recording steps below apply regardless of which capture tool you use to produce the file you upload. Follow them in order.
Get the technical setup right before you sit down. A clean frame with even light gives the model the visual signal it needs.
With the setup locked, the room around you and what you wear become the next things to control.
The model learns from everything in the frame, including reflections and background sound. Pick a setting that stays out of the way.
A clean environment sets up the next speaking segment.
Now record the speaking half of the clip. Tavus currently specifies a minimum of 30 seconds of speech within the roughly two-minute recording.
Once the speaking segment is captured, the still segment gives the model its listening baseline.
The still segment teaches the model how you look while someone else is speaking. Stay alone in frame throughout.
Clean footage from both segments produces a usable Replica; the next choice is what workflow it runs in.
A Custom Replica trained from your two minutes of video, or a Stock Replica pulled from the pre-built library, can run in almost any workflow that involves conversation. Teams configure the workflow without code in PAL Maker, or white-label it through the Conversational Video Interface (CVI) API. What changes across use cases is the knowledge, the guardrails, and the handoff logic.
Inbound qualification calls follow a predictable script until they don't. Configure the workflow so the clone handles the routine and escalates the exceptions.
Orum embedded Tavus-powered role-play, so reps practice cold calls, discovery, and objection handling on demand, with scenarios tuned by buyer type, difficulty, and sales stage. Managers report more confident performance on real calls, and 3X sales meetings.
Care conversations often turn on what a patient doesn't say out loud. Configure the workflow to gather information, explain plainly, and know when to hand off.
CareFlick, an AgeTech company, runs companion conversations for isolated seniors on this pattern, and achieves 54% user retention.
Practice conversations need to feel real enough that users see improvement. Configure the workflow so the clone challenges the user without breaking character.
Final Round AI scaled lifelike mock interviews for 100K+ users, logging 1.2M practice minutes with Tavus CVI.
Two minutes of video is a small ask. What separates a clone that gets used from one that sits in a demo is everything that happens after the upload: the knowledge it draws on, the guardrails around what it can say, and whether it can perceive the person in front of it and respond in the moment.
Tavus, the human computing company, builds PALs so one person's presence can reach conversations they would otherwise decline. The behavioral stack listens, holds the floor, decides what to say, and renders the response as a face-to-face exchange rather than a talking-head playback.
See it for yourself. Book a demo.
Custom Replicas are trained from about two minutes of recorded video. Processing runs after upload; check the Tavus training docs for the current turnaround before you plan a launch date.
Two minutes is enough for current training paths to capture facial geometry, skin tone, texture, listening posture, and person-specific voice detail. Recording quality, lighting, and the split between speaking and still segments affect the outcome more than raw length.
It depends on the system behind the face. A static clone renders finished video from typed text; a PAL runs on real-time infrastructure that perceives audio and visual cues, manages turn-taking, and renders responsive facial behavior while the other person is still on the call.
Explicit, documented consent from the person being cloned. EU Article 50 has required disclosure of AI-generated or manipulated likenesses since 2 August 2026, and other jurisdictions have parallel rules; check local law before deployment.
A Custom Replica is trained from two minutes of a specific person's video and reproduces that individual. A Stock Replica comes ready-made from a pre-built library and requires no recording.