The people an organization most wants in a conversation often have the least room on their calendar. A founder closes every deal she joins. A clinician earns trust in a single visit. Their value shows up in conversation, and each conversation takes time.

An AI clone learns how a specific person looks and sounds from about two minutes of video, then carries that presence into more conversations than a calendar could ever hold.

It can hold a live exchange or read a script aloud. The result depends on decisions made before the camera turns on, on which category of system receives the footage, and on the consent that separates a legitimate clone from impersonation.

What an AI clone actually is

An AI clone is a software model of one specific person, built from a short clip of video and audio. It reproduces a person's face, voice, and behavioral patterns, and once trained, it can say things its subject never recorded.

Consent separates a legitimate clone from impersonation. EU Article 50 has required disclosure of AI-generated or manipulated likenesses since 2 August 2026.

Tavus, the human computing company, builds Personified Application Layers (PALs) for live exchange. A PAL is a real-time application you talk to and build a relationship with, one that sees, hears, remembers, and responds face-to-face across text, voice, and video, rather than a chatbot you ping for one-off answers. That live presence starts with what the model can learn from the training footage.

What a real-time PAL adds to a clone

A clone can do two things with a script: play it back as a finished video, or use it as raw material for a live conversation. Playback works for an announcement or a training module. The live option is what most enterprise teams want, because it produces outcomes a static clip cannot.

A PAL delivers that live presence through a closed loop of models working together, and the practical wins show up in the conversation itself:

  • It reads the full signal. Raven-1 fuses tone, expression, hesitation, and gaze into a single read of how the other person actually feels, catching a polite objection dressed as agreement.  
  • It knows when to speak and when to wait. Sparrow-2 predicts who owns the conversational floor from streaming audio, so the clone holds space through a thinking pause and stops mid-clause when someone cuts in.  
  • It answers with the right information. The LLM layer, hosted or bring-your-own, decides what to say and pulls verified detail from the Knowledge Base in about 30ms.  
  • It looks like it is listening. Phoenix-4.5 renders responsive facial behavior at 40 fps in 1080p, with the nods and micro-expressions that tell a person on the other end they have been heard.  
  • It remembers across sessions. Persistent Memory carries context, preferences, and progress from one conversation to the next, so the second call picks up where the first left off.

Those capabilities are what a static clip cannot deliver, and what makes a two-minute recording worth training in the first place.

What a modern AI clone learns from two minutes of footage

Two minutes is enough for current Replica training paths to capture what makes one person recognizable on video. The clip pairs a speaking segment with a still segment, and each half teaches the model something specific.

From that recording, the model learns:

  • Facial geometry. The shape of the face, the placement of features, and the proportions that hold across expressions.  
  • Skin tone and texture. Color, contrast, and surface detail under the lighting used during capture.  
  • Speaking behavior. How the mouth moves, how the head shifts, and how expressions form during natural speech.  
  • A listening baseline. From the still segment, the resting posture and gaze the model returns to while the other person is speaking.  
  • Voice. Person-specific pitch, cadence, and timbre from the audio, though recording quality and length shape the result.

Capturing all of this cleanly depends on how you set up the recording and what the subject does on camera.

How to record the footage that trains your clone

In PAL Maker, the no-code setup flow walks you through uploading the clip, selecting a training path, and previewing the resulting Replica besfore it goes live. The recording steps below apply regardless of which capture tool you use to produce the file you upload. Follow them in order.

1. Set up your camera and lighting

Get the technical setup right before you sit down. A clean frame with even light gives the model the visual signal it needs.

  • Place the camera at eye level, with your head, shoulders, and upper chest in frame, and your face filling at least a quarter of it.  
  • Use soft, diffuse light so no shadows fall across your face, and add a backlight for dark hair so it doesn't vanish into the backdrop.  
  • Record at 1080p and 25 frames per second (fps) or better in a desktop app such as QuickTime or Windows Camera, not a browser recorder.

With the setup locked, the room around you and what you wear become the next things to control.

2. Control your environment and wardrobe

The model learns from everything in the frame, including reflections and background sound. Pick a setting that stays out of the way.

  • Choose a neutral background your clothes don't blend into, and record in a quiet room since voice cloning reproduces background noise.  
  • Skip patterned or shiny clothing, which can create moiré and glare on camera.  
  • Remove glasses, jewelry, over-ear headphones, and anything else that covers your face or neck.

A clean environment sets up the next speaking segment.

3. Speak naturally for the required segment

Now record the speaking half of the clip. Tavus currently specifies a minimum of 30 seconds of speech within the roughly two-minute recording.

  • Talk at your normal conversational pace, keeping your eyes on the lens.  
  • Keep head and body movement subtle, with natural pauses between thoughts.  
  • Smile widely for at least two seconds at some point during the segment.

Once the speaking segment is captured, the still segment gives the model its listening baseline.

4. Hold still for the listening segment

The still segment teaches the model how you look while someone else is speaking. Stay alone in frame throughout.

  • Hold still for at least 30 seconds with your lips closed and neutral, gaze on the camera, and no head tilting.  
  • Keep the same lighting and framing you used for the speaking segment.  
  • Avoid glancing off-camera to read notes, since angled frames rarely appear in training data.

Clean footage from both segments produces a usable Replica; the next choice is what workflow it runs in.

How to configure your clone for different workflows

A Custom Replica trained from your two minutes of video, or a Stock Replica pulled from the pre-built library, can run in almost any workflow that involves conversation. Teams configure the workflow without code in PAL Maker, or white-label it through the Conversational Video Interface (CVI) API. What changes across use cases is the knowledge, the guardrails, and the handoff logic.

Sales qualification and outreach

Inbound qualification calls follow a predictable script until they don't. Configure the workflow so the clone handles the routine and escalates the exceptions.

  • Connect the current rate card, product catalog, and objection responses through the Knowledge Base to keep answers accurate.  
  • Set an Objectives and Guardrails rule to confirm budget and timeline before the clone offers a meeting.  
  • Route qualified calls to a live calendar and log the transcript for the account owner.

Orum embedded Tavus-powered role-play, so reps practice cold calls, discovery, and objection handling on demand, with scenarios tuned by buyer type, difficulty, and sales stage. Managers report more confident performance on real calls, and 3X sales meetings.

Patient intake and follow-up care

Care conversations often turn on what a patient doesn't say out loud. Configure the workflow to gather information, explain plainly, and know when to hand off.

  • Load intake questions, medication schedules, and post-visit education material into the Knowledge Base.  
  • Set a medical guardrail such as "don't give medical advice" to route clinical questions to a nurse.  
  • Enable Health Insurance Portability and Accountability Act (HIPAA) settings on an eligible Enterprise plan and sign a business associate agreement before any patient data enters a conversation.

CareFlick, an AgeTech company, runs companion conversations for isolated seniors on this pattern, and achieves 54% user retention.

Candidate screening and employee coaching

Practice conversations need to feel real enough that users see improvement. Configure the workflow so the clone challenges the user without breaking character.

  • Load interview questions, role rubrics, or sales playbooks into the Knowledge Base for reference.  
  • Use Objectives to define a completion criterion, such as covering three qualification questions in a discovery role-play.  
  • Enable Function Calling to log scores, send summaries, or trigger a follow-up session.

Final Round AI scaled lifelike mock interviews for 100K+ users, logging 1.2M practice minutes with Tavus CVI.

The footage is the easy part

Two minutes of video is a small ask. What separates a clone that gets used from one that sits in a demo is everything that happens after the upload: the knowledge it draws on, the guardrails around what it can say, and whether it can perceive the person in front of it and respond in the moment.

Tavus, the human computing company, builds PALs so one person's presence can reach conversations they would otherwise decline. The behavioral stack listens, holds the floor, decides what to say, and renders the response as a face-to-face exchange rather than a talking-head playback.

See it for yourself. Book a demo.