Matched mouth movements, word by word

AI Lip Sync

Give ClipNova a face and the words: your own recording or a typed script. It animates the face so the lips, head and expression follow the speech. Already filmed? Lip-sync your video into another language.

Perfect for
Trusted by creators
worldwide
Picture of AI Lip Sync
How it works

A face and a voice in. A talking video out.

The first three steps run in the AI Talking Avatar tool, the fourth in the AI Video Translator. Here is the same presenter followed from a still photo to a lip-synced video.

  1. Picture of AI Lip Sync — a front-facing portrait of a presenter chosen as the face to animate
    — 01 —
    The face

    Start from a photo of a face

    Pick an avatar from the library, generate a new portrait from a description, or upload your own photo. A clear, front-facing face in good light gives the most natural mouth movement.

  2. Picture of AI Lip Sync — an uploaded voice recording and a typed script, the two ways to give the face its words
    — 02 —
    The audio

    Upload a recording, or type a script

    Upload a speech clip and the face says exactly that, in your voice. Or type a script and pick a ClipNova voice to read it. You can also place the person in a new background and choose a 16:9, 9:16 or 1:1 format.

  3. Picture of AI Lip Sync — the presenter now talking in a video player, with the speech waveform under the frame
    — 03 —
    The sync

    The lips, head and expression follow the speech

    Kling AI Avatar v2 (Base and Pro) or OmniHuman (Ultra) animates the photo to the audio: the mouth matches each sound, and the head and face move the way a speaker's do. Longer speech is split into lip-synced segments and joined into one video.

  4. Picture of AI Lip Sync — an original English video frame next to the same frame lip-synced in Spanish
    — 04 —
    Already filmed?

    Lip-sync your video into another language

    Upload a 2 to 60 second talking video to the AI Video Translator. It is dubbed into one of 18 languages in the speaker's own voice, and lipsync-2 redraws the mouth on your original footage so it matches the new words.

Why it looks real

Lip sync is more than moving a mouth.

A lip-sync that only moves the jaw looks wrong within a second. People read faces: the shape of the lips on an M or an O, a nod at the end of a sentence, the eyes that keep looking at them. ClipNova uses talking-avatar models that animate the whole face to the audio, not just the mouth.

You choose how the words arrive. Upload your own recording when the delivery matters, or type a script and let a ClipNova voice read it. Either way the timing comes from the audio, so the lips land on every syllable.

For footage you already have, the AI Video Translator does the reverse: it keeps your picture, dubs the speech into another language in the speaker's own voice and re-syncs the mouth to it. No re-shoot, and the video still looks like your video.

Built for

Any face, any words

Ads and UGC

Spokesperson videos without a shoot

Turn a product script into a presenter talking to camera in 9:16, ready for Reels and TikTok.

Creators

A face for a faceless channel

Give your narration a consistent on-screen host without filming yourself.

Training and support

Explainers with a human face

Turn help-center scripts and onboarding steps into short talking videos.

Localisation

The same video in a new language

Re-sync an existing talking video to a translated dub that keeps the speaker's voice.

AI lip sync questions.

What is AI lip sync?+

AI lip sync matches a face's mouth movements to a piece of speech. In ClipNova you can animate a photo so it speaks your audio or script, or re-sync a talking video to a translation of what was said.

Can I lip-sync a photo to my own audio?+

Yes. In the AI Talking Avatar tool, choose or upload a face, then upload your speech clip instead of typing a script. The video follows your recording word for word.

Can I lip-sync a video I already filmed?+

Yes, into another language. The AI Video Translator dubs a 2 to 60 second talking video into one of 18 languages in the speaker's own voice, then lip-syncs the mouth on your original footage to the new words.

Which models are used?+

Talking videos use Kling AI Avatar v2 on the Base and Pro tiers and OmniHuman on Ultra. Video translation uses ElevenLabs dubbing followed by the lipsync-2 model.

What makes a good source photo or video?+

One person, face visible and turned toward the camera, evenly lit and sharp. Clean speech with no music on top gives the most accurate mouth shapes.

Which formats can I export?+

Talking videos can be made in 16:9, 9:16 or 1:1, or in the photo's own shape. Translated videos keep the format of the clip you uploaded.

How much does it cost?+

Lip-synced video is included in every paid plan, from $19 a month, and every plan comes with full commercial rights. The cost of each video is shown in credits before you generate.

Start creating today

A face and a voice are all it takes.

Join creators turning ideas into scroll-stopping videos, no crew, no software, no learning curve.

Make a lip-sync video