ElevenLabs launches v4 speech models with faster cloning and 90 language support

ElevenLabs released v4 and v4 Turbo speech models with 10-second voice cloning and support for 90 languages as its revenue run rate surpasses $600 million. The v4 Turbo variant cuts conversational latency by generating audio as soon as the language model starts answering.

ElevenLabs launches v4 speech models with faster cloning and 90 language support

ElevenLabs released two new speech models Monday - ElevenLabs v4 and v4 Turbo - that bring more granular expression control, faster voice cloning, and lower latency for conversational voice agents. The launch matters for teams in marketing, sales, customer support, and content production who rely on voice output at scale, as the company now counts more than 55% of its business from large enterprises and has seen its annualized revenue run rate climb past $600 million.

The v4 generation adopts a new architecture that the company says allows for better control and quicker cloning. Users can now clone a voice with just 10 seconds of audio. The model also handles voice identity more reliably over longer stretches of text and reads with an awareness of context, adjusting expression as the text unfolds.

Expression control and language expansion

ElevenLabs introduced inline expression tags with its v3 model and is extending that system in v4. Users can now stack multiple tags and the model will follow the sequence - a shift that gives voiceover artists, marketers, and content teams finer control over delivery without manual retakes. The previous version supported 70 languages. The new release pushes that number to 90, with the biggest quality gains observed in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.

Lower latency for voice agents

The v4 Turbo variant is built for speed in two-way conversations. ElevenLabs said the model can start generating audio as soon as the underlying large language model begins producing answers, cutting the pause that often makes automated calls feel stilted. The model also handles confrontations, escalations, and hold patterns differently, which the company frames as a step toward better issue resolution in customer support and hospitality settings.

The startup's enterprise calling business has scaled quickly. Co-founder and CEO Mati Staniszewski told TechCrunch the company is aiming for an IPO "in the next years," though he did not commit to a timeline. ElevenLabs raised $500 million from Sequoia earlier this year at an $11 billion valuation, and rumors of a follow-up round at $22 billion are already circulating. Headcount has passed 800 as the company hires across India, Europe, and Brazil.

Why this matters for creatives, support teams, and communicators

For anyone producing voice content - whether that is a marketing team localizing ads into Portuguese and Mandarin, a support leader evaluating conversational AI, or a writer turning scripts into audio - the v4 release changes two practical constraints. First, the 10-second cloning threshold lowers the bar for creating consistent voice personas without lengthy studio sessions. Second, the lower latency and interruption handling in v4 Turbo make it feasible to deploy voice agents in frontline roles where a half-second delay can break trust. Professionals exploring these capabilities can build the relevant skills through Text to Speech AI Courses and AI Voice Modulation Courses.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)