# Mivo Sync > AI lip sync video generation for dubbing, talking avatars, and multilingual localization. Upload a video or photo, add audio or text, and generate a natural lip-synced MP4 in minutes. Start with 20 free credits, no watermark, and 30+ languages. ## About Mivo Sync AI Lip Sync Video Generator for Photos, Videos & Scripts Upload a video or photo, add audio or text, and create a natural lip-synced MP4 in under 2 minutes. No editing skills required. ### Use cases - **Turn one creator video into local channels**: Keep the same face and delivery while replacing the voice track for each market. Build Shorts, TikTok, and YouTube versions without another shoot. - **Produce UGC ad variants without new talent**: Reuse a proven spokesperson, swap the script or language, and generate market-specific ads for Shopify, TikTok Shop, and paid social tests. - **Dub vertical drama episodes market by market**: Sync translated dialogue back onto the original cast so every episode feels native in each region without reshooting scenes. - **Go from upload to review in minutes**: Use short-turnaround renders for news clips, product posts, training updates, and daily social publishing queues. - **Sell product videos around the world**: Export teams turn one product demo, factory intro, or buyer testimonial into 30+ language versions for global buyers, storefronts, and overseas social channels. ### How it works 1. **Upload Your Video or Photo**: Add any front-facing video clip or portrait photo. MP4 and MOV are supported, and still photos can become talking avatars. 2. **Add Your Audio or Script**: Upload an audio file, record directly, or type a script and choose an AI voice. The workflow supports English, Spanish, French, German, Hindi, and more. 3. **Generate and Download**: The AI maps mouth shapes frame by frame and renders a lip-synced MP4. Download watermark-free or share directly from the result screen. ## Create lip-synced videos with an AI Talking Photo Generator Turn one portrait and either recorded audio or a written script into a talking video. Choose a model, generate the lip-synced result, and download the finished MP4. ### Overview Choose a model, add the portrait and speech source, and the form checks the model-specific requirements before generation. Unlike a general video generator or stock-avatar tool, this workflow animates the visible subject from your image to follow your chosen speech. ### Use cases - **Podcast repurposing**: Pair a host portrait with an approved excerpt from a longer episode to create a shorter face-led video for social publishing. - **Founder-led marketing**: Pair a founder voice memo or written script with a clean portrait for product updates, launch messages, and sales follow-ups. - **Multilingual localization**: Keep one approved brand portrait and create regional clips from translated audio or scripts for announcements, help content, and campaign variants. ### How to use 1. **Upload the portrait**: Use a clear JPEG, PNG, or WebP portrait with the face visible and unobstructed. The form checks the selected model's image requirements before submission. 2. **Add audio or a script**: Upload an MP3, WAV, M4A, or AAC voice track, or write a script and generate speech inside the form. Use one speech source for each request. 3. **Render the MP4**: Review the displayed model and credit cost, generate the talking photo, then preview or download the successful MP4 from your saved result. ### Features - **Use portraits and voices you are allowed to use**: Only upload a portrait, recording, or script when you own it or have the necessary permission. Do not create deceptive impersonations or content that misrepresents the speaker. - **Input limits depend on the selected model**: Supported image dimensions, file size, audio duration, and credit cost can vary by model. The form validates the active selection and shows the relevant requirements before submission. - **Temporary AI inputs follow a cleanup lifecycle**: Confirmed temporary portrait and audio inputs enter Mivo Sync's seven-day cleanup lifecycle. This cleanup policy is separate from a successful generated result saved to the account. - **Successful videos remain in Assets**: A successful talking photo is stored in your account's asset library as an MP4 result. You can preview or download it and return to the saved asset later. ### Testimonials > I can pair an approved portrait with a revised launch script and create a fresh spokesperson clip without arranging another shoot. > — Luke Lawrence, Brand designer > Talking-photo drafts make it easier to test which message deserves a full production pass before our team invests in filming. > — Nihal Okur, E-commerce marketer > I reuse one approved brand portrait for short social updates, then switch the audio or script for each campaign. > — Mélina Clement, Social media manager ### FAQ **What is an AI talking photo generator?** It creates a video from one still portrait and a speech source. Upload a portrait, then provide recorded audio or write a script for generated speech. The tool animates the visible face to follow that speech and returns a downloadable MP4 when the task succeeds. **What kind of photo works best for a talking photo?** Use one clear portrait with the face and mouth visible, even lighting, and little obstruction. Avoid masks, heavy shadows, very small faces, or extreme side angles. Exact dimensions and file-size limits depend on the selected model, and the form checks them before submission. **How is a talking photo different from a stock avatar or general video generator?** This workflow starts from the portrait you upload and a specific audio track or script. A stock-avatar tool begins with a preset character, while a general video generator may create an entire scene. A talking photo focuses on producing a face-led video from your chosen portrait and speech source. **What photo and audio formats are supported?** Upload a portrait as JPEG, PNG, or WebP and recorded speech as MP3, WAV, M4A, or AAC, or enter a script instead of an audio file. Size, duration, and image-dimension requirements vary by model and are shown in the form. Successful video results are provided as MP4 files. **How are my portrait, audio, and generated video handled?** Only use inputs that you own or have permission to use. Confirmed temporary AI inputs enter Mivo Sync's seven-day cleanup lifecycle. A successful generated video is stored in your account's asset library so you can preview or download it later. **How many credits does a talking photo use?** The required credits depend on the selected model and the speech duration. The form calculates and displays the cost before submission. Available signup or promotional credits vary by account and market, so the page does not promise a fixed free balance. ## AI Voice Cloning from a short, authorized sample Upload 3–30 seconds of clean speech, type new words, and generate audio that follows the authorized speaker's vocal character. The same reference can speak in 10 supported output languages. ### Overview Unlike a general AI voice generator that starts from a preset or designed voice, voice cloning starts from a specific speaker sample. It also differs from a voice changer, which transforms an existing recording, and ordinary text-to-speech, which does not try to reproduce that speaker's identity. ### Use cases - **Creator corrections**: Fix a name, date, instruction, or call to action without rebuilding the microphone setup. Paste the corrected sentence and export a replacement audio clip for the edit. - **Course and training updates**: Replace an outdated instruction without asking the instructor to re-record a full module. Generate the changed line, review it, and place it in the existing lesson timeline. - **Podcast intros and ad reads**: Prepare an approved intro, correction, sponsor line, or closing message from a host reference without scheduling another recording session. - **Multilingual marketing**: Keep the same recognizable speaker across English, Chinese, Japanese, Korean, French, German, Italian, Spanish, Portuguese, or Russian campaign scripts. ### How to use 1. **Upload the reference and confirm permission**: Choose 3–30 seconds of clean speech from a voice you own or have explicit permission to use. Keep the file under 20 MB with one clear speaker and little music, echo, or background noise. 2. **Enter the script and delivery settings**: Enter up to 2,000 characters, select a language or use automatic detection, and optionally add the reference transcript and a style instruction of up to 500 characters. 3. **Generate, review, and save the result**: Review the displayed credit cost, submit the request, and listen to the result. A successful audio file can be downloaded from the result panel and remains available in your asset library. ### Features - **Use only an authorized voice**: Every request requires confirmation that you own the voice or have explicit permission from the speaker. Do not use cloned speech for impersonation, fraud, harassment, or deceptive content. - **Keep the reference within the input limits**: Use 3–30 seconds of speech in an accepted audio format, keep the upload under 20 MB, and choose a clip with one clear speaker. Source quality directly affects the generated result. - **Temporary inputs follow a cleanup lifecycle**: Confirmed temporary AI inputs enter Mivo Sync's seven-day cleanup lifecycle. This applies to the reference material used to run the task, not to a successful result saved in your account. - **Successful audio remains in Assets**: When generation succeeds, the audio result is stored in your account's asset library. You can preview or download it from the result panel and return to the saved asset later. ### Testimonials > I use the clone to patch a sentence after a lesson is recorded. Matching the original delivery keeps me from reopening the whole recording setup. > — Janick Mercier, Course creator > For podcast corrections, a short clean reference clip gives me a practical way to replace one line without rerecording the complete segment. > — Cecilia Paredes, Podcast producer > The language controls help our training updates keep one approved voice across regional versions, while the authorization step keeps the workflow clear. > — Lija Nogueira, Localization manager ### FAQ **What is AI voice cloning and how does it create new speech?** AI voice cloning analyzes a reference recording to capture recognizable speaker traits such as tone, rhythm, accent, and delivery. It combines those traits with new text to synthesize a fresh audio file. Mivo Sync performs task-by-task generation, so you provide a reference clip and script for each request rather than enrolling a permanent voice profile. **How do I clone my voice online with Mivo Sync?** Upload 3-30 seconds of clean speech, enter the new script, select an output language, and confirm that you own the voice or have permission to use it. A reference transcript and style note are optional. Submit the task, wait for the result panel to update, then preview or download the generated audio. **How much reference audio gives the clearest cloned voice?** Three seconds is the technical minimum, but clarity matters more than filling the 30-second limit. Use one speaker at a natural pace with little echo, music, or background noise. A clean sentence that includes normal pitch changes usually gives the model more useful speaker information than a longer noisy recording. **Which audio formats and output languages does voice cloning support?** Reference uploads can be MP3, WAV, M4A/AAC, OGG, or WebM, up to 20 MB and 30 seconds. Output text can be detected automatically or set to Chinese, English, Japanese, Korean, French, German, Italian, Spanish, Portuguese, or Russian. Cross-language results still depend on clear source speech and well-formed target text. **How is this different from persistent professional voice cloning?** Mivo Sync is designed for instant, task-by-task generation from a short sample. Professional voice-cloning services may train a reusable voice identity from much longer recordings and add team-level voice management. Use Mivo Sync for quick pickups, localized lines, and creator audio; use a persistent voice program when long-form consistency and managed voice enrollment are the primary requirement. **Can I clone another person's voice with this tool?** Only when that person has given you explicit permission. The form requires you to confirm ownership or authorization before every generation, and the API rejects requests without that confirmation. Do not use cloned speech for impersonation, fraud, harassment, or audio that hides who actually created it. **What happens to my reference audio and generated result?** The server accepts only reference uploads owned by the signed-in account and validates the file before creating a task. Confirmed temporary AI inputs enter the seven-day cleanup lifecycle. Successful generated audio is stored in your account's asset library so you can return to the result after leaving the page. **How many credits does AI voice cloning use?** A request uses a minimum of 5 credits. Once length-based pricing exceeds that minimum, the total is 2 credits for each started block of 100 characters across the script, reference transcript, and style instruction. The form shows the required credits before submission; available signup or promotional credits vary by account and market.