Give ClipNova a face and the words: your own recording or a typed script. It animates the face so the lips, head and expression follow the speech. Already filmed? Lip-sync your video into another language.
The first three steps run in the AI Talking Avatar tool, the fourth in the AI Video Translator. Here is the same presenter followed from a still photo to a lip-synced video.
Pick an avatar from the library, generate a new portrait from a description, or upload your own photo. A clear, front-facing face in good light gives the most natural mouth movement.
Upload a speech clip and the face says exactly that, in your voice. Or type a script and pick a ClipNova voice to read it. You can also place the person in a new background and choose a 16:9, 9:16 or 1:1 format.
Kling AI Avatar v2 (Base and Pro) or OmniHuman (Ultra) animates the photo to the audio: the mouth matches each sound, and the head and face move the way a speaker's do. Longer speech is split into lip-synced segments and joined into one video.
Upload a 2 to 60 second talking video to the AI Video Translator. It is dubbed into one of 18 languages in the speaker's own voice, and lipsync-2 redraws the mouth on your original footage so it matches the new words.
A lip-sync that only moves the jaw looks wrong within a second. People read faces: the shape of the lips on an M or an O, a nod at the end of a sentence, the eyes that keep looking at them. ClipNova uses talking-avatar models that animate the whole face to the audio, not just the mouth.
You choose how the words arrive. Upload your own recording when the delivery matters, or type a script and let a ClipNova voice read it. Either way the timing comes from the audio, so the lips land on every syllable.
For footage you already have, the AI Video Translator does the reverse: it keeps your picture, dubs the speech into another language in the speaker's own voice and re-syncs the mouth to it. No re-shoot, and the video still looks like your video.
Turn a product script into a presenter talking to camera in 9:16, ready for Reels and TikTok.
Give your narration a consistent on-screen host without filming yourself.
Turn help-center scripts and onboarding steps into short talking videos.
Re-sync an existing talking video to a translated dub that keeps the speaker's voice.
AI lip sync matches a face's mouth movements to a piece of speech. In ClipNova you can animate a photo so it speaks your audio or script, or re-sync a talking video to a translation of what was said.
Yes. In the AI Talking Avatar tool, choose or upload a face, then upload your speech clip instead of typing a script. The video follows your recording word for word.
Yes, into another language. The AI Video Translator dubs a 2 to 60 second talking video into one of 18 languages in the speaker's own voice, then lip-syncs the mouth on your original footage to the new words.
Talking videos use Kling AI Avatar v2 on the Base and Pro tiers and OmniHuman on Ultra. Video translation uses ElevenLabs dubbing followed by the lipsync-2 model.
One person, face visible and turned toward the camera, evenly lit and sharp. Clean speech with no music on top gives the most accurate mouth shapes.
Talking videos can be made in 16:9, 9:16 or 1:1, or in the photo's own shape. Translated videos keep the format of the clip you uploaded.
Lip-synced video is included in every paid plan, from $19 a month, and every plan comes with full commercial rights. The cost of each video is shown in credits before you generate.
Join creators turning ideas into scroll-stopping videos, no crew, no software, no learning curve.