The same person, speaking another language

AI Video Translator

Upload a talking video and pick a language. ClipNova translates the speech, re-voices each speaker in their own voice, keeps the music and background sound, then lip-syncs the mouth to the new words.

Perfect for
Trusted by creators
worldwide
Picture of AI Video Translator
How it works

Upload once. Speak to a new market.

Here is the same clip followed through every step: a chef filmed in English, translated into Spanish with her own voice and matching lips.

  1. Picture of AI Video Translator — a chef's talking video uploaded in a player, 42 seconds long
    — 01 —
    Your video

    Upload a 2 to 60 second talking video

    MP4, MOV or WEBM, up to 200 MB. One person facing the camera in good light gives the cleanest result, and clips with several speakers work too.

  2. Picture of AI Video Translator — the language picker set to translate from English into Spanish
    — 02 —
    Language

    Pick the language to translate into

    18 languages are available: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Korean, Mandarin Chinese, Arabic, Hindi, Indonesian, Dutch, Polish, Turkish, Vietnamese and Swedish. The credits for the clip are reserved up front, and whatever the run does not use comes back to you.

  3. Picture of AI Video Translator — the chef's clip with two audio tracks, her own voice in Spanish and the original background music
    — 03 —
    Dub

    Dubbed in the speaker's own voice

    ElevenLabs dubbing separates the speech from the rest of the soundtrack, translates it and re-voices each speaker with a clone of their own voice, on the original timing. Then it mixes the music and background sound back in, so nothing else about the clip changes.

  4. Picture of AI Video Translator — the original English frame next to the translated Spanish frame with lip-synced speech
    — 04 —
    Lip-sync

    The lips follow the new words

    A lip-sync model (lipsync-2) redraws the mouth on the original picture so it matches the translated speech. You get one MP4 to download, with the same framing, the same voice and the new language.

Why it feels native

Subtitles get read. A dubbed voice gets watched.

Most video translators stop at subtitles, or replace the speaker with a stock voice that sounds nothing like them. ClipNova's AI Video Translator keeps the person: each speaker is re-voiced in a clone of their own voice, on the same timing as the original take.

The rest of the soundtrack is kept as well. Music, room tone and sound effects are separated from the speech before the dub and mixed back in after it, so a translated clip still sounds like the clip you filmed.

The last step is what makes a dub believable: the mouth is lip-synced to the translated words on your original footage. No re-shoot, no voice actor, no editor, and one upload gives you a version for every market you want to reach.

Built for

One video, every language you sell in

Creators

Reach a second audience

Publish your best Reels, Shorts and TikToks in Spanish, Portuguese or Hindi in your own voice.

Brands and ads

Localise spokesperson ads

Run the same UGC or founder ad in every market without filming it again.

Courses and training

Multilingual lessons

Translate short lesson clips and onboarding videos while the instructor keeps their voice.

Interviews and podcasts

Several speakers, one dub

Clips with more than one speaker are dubbed with each person in their own voice.

AI video translator questions.

What does an AI video translator do?+

It turns a video recorded in one language into the same video in another language. ClipNova translates the speech, re-voices it in the speaker's own voice, keeps the background audio and lip-syncs the mouth to the new words.

Which languages can I translate into?+

18 languages: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Korean, Mandarin Chinese, Arabic, Hindi, Indonesian, Dutch, Polish, Turkish, Vietnamese and Swedish.

Does the speaker keep their own voice?+

Yes. There is no voice to pick: ElevenLabs dubbing re-voices each speaker with a clone of their own voice, so the translated video still sounds like them.

Is the background music kept?+

Yes. The speech is separated from the music and ambient sound, translated, and the background is mixed back in under the new voice.

How long can my video be?+

Between 2 and 60 seconds per clip, in MP4, MOV or WEBM, up to 200 MB. Translate a longer video in parts.

Does it work with several speakers?+

Yes. Each speaker is dubbed in their own voice. For the cleanest lip-sync, faces should be visible, well lit and turned toward the camera.

How much does it cost?+

Video translation is included in every paid plan, from $19 a month. Credits for the clip are reserved before the run and anything it does not use is returned to your balance.

Start creating today

Your video already works. Now make it speak their language.

Join creators turning ideas into scroll-stopping videos, no crew, no software, no learning curve.

Translate a video