D-ID Expands Video Translation Ecosystem with Multilingual Voice Cloning and Lip Sync
D-ID detailed its video translation toolset designed to automate multilingual dubbing and realistic lip synchronization for global video creators. The platform allows editors to translate talking-head clips into dozens of languages while matching speaker cadence and visual mouth movements.
D-ID, the AI video generation platform, has highlighted its video translation and multilingual dubbing capabilities tailored for digital creators and enterprise localization workflows. As of early 2026, the tool combines automatic audio translation, voice cloning, and visual lip-sync adjustment into a single pipeline. The workflow allows creators to take existing talking-head footage and re-dub it into different languages while preserving the original speaker's vocal characteristics and updating facial movements to match translated phonemes.
What's new
D-ID's video translation architecture processes raw video files by extracting spoken dialogue, converting speech to text, and running neural machine translation across more than 30 target languages. The system then synthesizes cloned audio matching the original tone, pitch, and timing before applying automated facial animation updates to realign lip movements with the new audio track.
The translation features handle both AI-generated avatars created on the platform and uploaded real-human video footage. Users can review translated scripts prior to rendering, edit translated text for localized accuracy, and select specific voice parameters to adjust speed or accent nuances before generating final video files.
How it fits your workflow
For video editors, educators, and global marketing teams, D-ID eliminates the traditional multi-step translation pipeline that required separate transcription services, voice actors, and manual re-editing. Instead of reshooting footage or relying on generic voiceovers with mismatched lip movements, editors upload a source MP4 or MOV file, pick their target output languages, and receive fully dubbed clips with synchronized mouth motion.
In the AI dubbing landscape, D-ID competes directly with specialized video translation tools like HeyGen, ElevenLabs Video Dubbing, Synthesia, and Rask AI. While ElevenLabs focuses primarily on audio quality and vocal expression matching, and HeyGen offers comparable visual re-dubbing, D-ID integrates translation directly into its existing synthetic avatar generation suite. This allows teams to create original avatar videos and localize existing footage within the same asset management interface.
What it costs / how to try it
D-ID offers video translation features through its standard web interface and API. Free trial accounts include initial generation credits for testing translation quality, with paid plans starting on monthly subscription tiers that scale based on video generation minutes and export resolution options.
Read the original announcement on D-ID ↗