07/14/2026
ElevenLabs has spent years building the most realistic AI voice on the market. Over 11,000 voices, support for 32+ languages, and voice cloning from roughly a minute of source audio.
They have now added a face to it.
Avatars generates photorealistic, lip-synced talking-head video from a photo or a text prompt. Because ElevenLabs owns the speech layer, the voice model and the lip-sync model run in the same environment, so audio and video are produced together in a single step rather than stitched across three or four tools.
For teams producing UGC-style content:
Performance marketing: ad variants at scale
Organic social: on-brand content without on-camera talent
Localization: one avatar, 30+ languages
Learning and development: onboarding and training
Product marketing: localized explainers
The workflow:
Navigate to Image & Video, open the Avatar section, and select New.
Upload 3 to 5 reference images of the same person from different angles, or describe the avatar using a text prompt.
Name the avatar and assign a default voice. Avatars are persistent identities, so they can be reused across unlimited videos.
Create Styles to vary outfit, background, lighting, and framing while preserving the same identity.
Select Create Lip Sync, pair the avatar with a voice from your library (including cloned voices), input your script, and generate.
To scale, add the Avatar node to a Flow and batch-execute the pipeline across products, languages, and hooks.
Worth noting before you plan around it: Avatars is available on paid plans only, generation runs on the standard Image & Video credit structure, and certain avatar models and reference-image uploads are currently restricted in the United States.
Save this as a reference for your next content sprint.