How to lip sync a video to new audio
Lip sync starts from audio, not from a script. If you already have the recording — a real voiceover, a dub in another language, an old take with a fixed script error — this matches the mouth to it instead of synthesising new speech. Use AI Actor instead if you only have text and need the voice generated too.
Open the AI Lip SyncStep by step
- 1
Upload the audio
Up to 120 seconds. The output video ends up exactly as long as this track, so trim it to the part you need first.
- 2
Provide a face — a portrait or a clip
A still portrait gives a static frame with just the mouth moving. A video clip keeps the original head movement and only the mouth is re-matched.
- 3
Set the real audio length
This drives the price, so enter it accurately.
What makes the difference
- Use a clip instead of a still portrait when you want the person's original gestures and head movement kept — only the mouth changes.
- Clean, well-recorded audio with clear speech gets the tightest sync. Music underneath the voice or heavy background noise makes the match less precise.
- This is a redub tool, not a translation tool — generate or source the new-language audio yourself first, then sync the mouth to it.
Questions
Can I use this to dub a video into another language?
Yes — generate or record the translated audio first with a tool like text to speech, then run it through here against the original footage.
What is the difference between this and AI Actor?
AI Actor starts from a script and synthesises the voice for you. Lip Sync starts from audio you already have and matches the mouth to it.
How much does it cost?
20 credits per second of audio. A 30-second track is 600 credits.