What Is ElevenLabs v4? The Emotive Text-to-Speech Model
5 min read

ElevenLabs released the fourth version of its text-to-speech (TTS) model, Eleven v4, along with a fast variant for real-time use, v4 Turbo. The company presents v4 as "our most emotive model yet." The new architecture reads a script the way a voice actor would: it knows who is speaking, what just happened and how each line should land.
This article summarizes the ElevenLabs v4 section of the AI Search channel's weekly AI news video on YouTube and extends it with the details on ElevenLabs' model page; the video link is at the end. This is a news roundup, and I have not tested the model. Information comes from ElevenLabs' own page; the "best model" ranking rests on a single leaderboard shown in the video.
What is ElevenLabs v4?
Eleven v4 is ElevenLabs' text-to-speech model for produced content (books, ads, video, games) where quality comes first. According to the page:
- 90+ languages, a wide emotional range, multiple speakers and built-in sound effects.
- New architecture: higher audio quality and a wider emotional range; speaker identity stays stable across regenerations, so redoing a line once or fifty times keeps the same person speaking.
- Context stitching: pacing and delivery stay consistent across long scripts; an audiobook sounds like a single take from start to finish.
- Voice cloning: Professional Voice Clones (PVC), unavailable in v3, are back in v4 and perform with the model's full emotional range. Instant Voice Clones (IVC) now outperform the Professional Voice Clones of Multilingual v2. Every clone requires verified consent from the voice's owner. Clones made before v4 need retraining to work well with it.
- Voice library: 17,500+ voices work with v4. You can also clone a voice from ten seconds of audio or design one from a one-sentence description.
Page: elevenlabs.io/v4.
How do you use metatags (audio tags)?
In v4, direction is given with natural-language tags in square brackets that you write into the script. Examples from the page:
- Pauses:
[pause],[long pause] - Emotion and delivery:
[whispers],[excited],[sighs],[warm],[amused],[nervous laugh] - Sound events:
[laughs],[door slams],[Gong sounds],[crowd applause]
In the demo on the page a game-show host and a contestant speak; the text looks like this: [warm] Welcome back to the final round. [long pause] Our returning champion needs one more answer. In the same demo, who is speaking is defined at the start of each line; the model follows tag sequences more reliably than v3.
An important note: SSML tags (such as <break>) are disabled in v4; use the natural-language tags above for pauses. Pronunciation dictionaries still work: you can define phonetic spellings for names, acronyms and technical terms (IPA is supported). A single generation takes up to 10,000 characters; for longer content, use context stitching.
A practical method: write the text without tags first, listen, then add tags only where you want emphasis. Sprinkling a tag on every line does not give natural results.
v4 Turbo: for real time and voice agents
v4 Turbo is the variant that delivers v4's emotional range at low latency. According to ElevenLabs, median inference latency is about 100 ms and time to first speech about 150 ms. In the company's own comparison chart, Cartesia Sonic 3.6 shows 262 ms and OpenAI GPT-4o mini TTS 814 ms. With bidirectional streaming (text in, audio out), audio starts before your language model finishes the sentence. It is used through the API and ElevenAgents, and works with the same Professional Voice Clones as v4.
Price and access
- Free plan: 10,000 credits a month (roughly 10 minutes of audio), personal use, no credit card. v4 uses the same credit pricing as other TTS models, so it is on every plan including the free tier.
- Paid plans: On the page they start at $6 a month; they include more credits, professional voice cloning and higher limits. The enterprise plan has custom pricing.
- Promotion: The page announces 3x credits on Creator and above until October 12.
- API: REST, streaming endpoints, TypeScript and Python SDKs. You choose the model with a single
model_id(eleven_v4). Output formats are MP3, WAV/PCM and µ-law for telephony.
The video's "10,000 credits on the free plan" is therefore incomplete: according to the page, it is a monthly allowance.
Security and privacy
Scripts and audio tags you send are not used to train models without your consent. ElevenLabs states SOC 2 Type II, ISO 27001 and PCI DSS Level 1 certifications, GDPR compliance and HIPAA-eligible workflows for healthcare. Enterprise customers can turn on zero retention mode. Generated audio can be detected as AI-generated by an AI speech classifier.
Who is it useful for?
- Video and game makers: character voices, narrator, trailers; sound design in a single pass with effect tags.
- Audiobook and podcast producers: context stitching and stable speaker identity.
- App developers: v4 Turbo for voice agents and real-time chat.
For generating game sound and music in code, see generating game music and sound with code; for small models that turn speech into text, see Whistle and Phonon 2.
Frequently Asked Questions
Is ElevenLabs v4 free?
It can be used on the free plan: 10,000 credits a month, roughly 10 minutes of audio, personal use. For more you need a paid plan.
What is a metatag?
A tag in square brackets written inside the text that sets the voice's tone, pause or sound effect; for example [whispers], [excited], [pause], [door slams].
What is the difference between v4 and v4 Turbo?
v4 is tuned for quality in produced content. v4 Turbo is the fast variant for voice agents and real-time use, offering the same emotional range at about 100 ms median inference latency.
Which languages does v4 support?
90+ languages. The page has no language list; check ElevenLabs' language list to see whether Turkish is included.
Is ElevenLabs v4 the best TTS model?
According to a leaderboard shown in the video, yes. But it is a single ranking; compare by listening with your own text.
Source
- AI Search, Gemini 4, GPT 6.1, Dots, Claude Sonnet 5.5, Ideogram 4.5, Flux 3: AI NEWS (YouTube, October 4, 2026): the ElevenLabs v4 section, examples and leaderboard.
- ElevenLabs, Eleven v4 and Eleven v4 Turbo: features, tags, price, API, privacy and FAQ.
Related Posts
What Is Comfy Agent? An Agent That Builds ComfyUI Nodes
Comfy Agent is an agent built into ComfyUI: describe what you want and it plans, builds and runs the workflow. Comfy Cloud only for now, free to try.
What Are InSpatio World 1.5, Sol Refiner and PixelUMM?
InSpatio World 1.5 roams a scene from an image, Nvidia's SoL-Refiner upscales to 4K in one step (27x faster on MiniMax H3), PixelUMM has no VAE.
Ideogram 4.5 vs Flux 3 Image: New AI Image Models
Ideogram 4.5 keeps pixels intact across multi-turn edits, Flux 3 Image adds 10 references and native 4K. Features, access and which fits which job.