Upload a talking video and pick a language. ClipNova translates the speech, re-voices each speaker in their own voice, keeps the music and background sound, then lip-syncs the mouth to the new words.
Here is the same clip followed through every step: a chef filmed in English, translated into Spanish with her own voice and matching lips.
MP4, MOV or WEBM, up to 200 MB. One person facing the camera in good light gives the cleanest result, and clips with several speakers work too.
18 languages are available: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Korean, Mandarin Chinese, Arabic, Hindi, Indonesian, Dutch, Polish, Turkish, Vietnamese and Swedish. The credits for the clip are reserved up front, and whatever the run does not use comes back to you.
ElevenLabs dubbing separates the speech from the rest of the soundtrack, translates it and re-voices each speaker with a clone of their own voice, on the original timing. Then it mixes the music and background sound back in, so nothing else about the clip changes.
A lip-sync model (lipsync-2) redraws the mouth on the original picture so it matches the translated speech. You get one MP4 to download, with the same framing, the same voice and the new language.
Most video translators stop at subtitles, or replace the speaker with a stock voice that sounds nothing like them. ClipNova's AI Video Translator keeps the person: each speaker is re-voiced in a clone of their own voice, on the same timing as the original take.
The rest of the soundtrack is kept as well. Music, room tone and sound effects are separated from the speech before the dub and mixed back in after it, so a translated clip still sounds like the clip you filmed.
The last step is what makes a dub believable: the mouth is lip-synced to the translated words on your original footage. No re-shoot, no voice actor, no editor, and one upload gives you a version for every market you want to reach.
Publish your best Reels, Shorts and TikToks in Spanish, Portuguese or Hindi in your own voice.
Run the same UGC or founder ad in every market without filming it again.
Translate short lesson clips and onboarding videos while the instructor keeps their voice.
Clips with more than one speaker are dubbed with each person in their own voice.
It turns a video recorded in one language into the same video in another language. ClipNova translates the speech, re-voices it in the speaker's own voice, keeps the background audio and lip-syncs the mouth to the new words.
18 languages: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Korean, Mandarin Chinese, Arabic, Hindi, Indonesian, Dutch, Polish, Turkish, Vietnamese and Swedish.
Yes. There is no voice to pick: ElevenLabs dubbing re-voices each speaker with a clone of their own voice, so the translated video still sounds like them.
Yes. The speech is separated from the music and ambient sound, translated, and the background is mixed back in under the new voice.
Between 2 and 60 seconds per clip, in MP4, MOV or WEBM, up to 200 MB. Translate a longer video in parts.
Yes. Each speaker is dubbed in their own voice. For the cleanest lip-sync, faces should be visible, well lit and turned toward the camera.
Video translation is included in every paid plan, from $19 a month. Credits for the clip are reserved before the run and anything it does not use is returned to your balance.
Join creators turning ideas into scroll-stopping videos, no crew, no software, no learning curve.