The gist
A voice actor is a separate line item: scheduling, a studio, an hourly rate. A speech synthesis platform is billed differently: you pay for volume (credits per month), and on long texts the difference is noticeable (current plans: What's current). Claude writes the script (a narration script or program code), ElevenLabs voices it, FFmpeg stitches the video together. No voice actor, no studio, no scheduling. Set up the pipeline (a chain of sequential steps) once, and from then on the factory runs by itself.
Key concepts
- TTS (Text-to-Speech): the technology for synthesizing speech from text
- Voice cloning: creating an AI copy of a voice, a digital clone of your voice from a sample
- ElevenLabs: one of the leading commercial TTS platforms
- ElevenLabs MCP (Model Context Protocol): a direct integration with Claude Code: a hosted server with OAuth sign-in, or a local server
- Eleven v4: ElevenLabs' new model (September 28, 2026): more control over intonation, 90+ languages; before it, Multilingual v2 was the main one
- Kokoro TTS: an open source alternative for offline voiceover (a limited set of languages, see below)
- TTS pipeline: Claude writes the script → TTS voices it → FFmpeg assembles the video
Theory
Why TTS belongs in an AI stack: a question of scale
Voiceover is the last manual step in content production. Claude writes text automatically, Midjourney generates images, FFmpeg assembles video. But the narrator is still a live person with a schedule and a price tag.
TTS closes that gap.
Three scenarios where TTS changes the economics:
Scenario 1: A YouTube channel at scale
One narrator can voice 4-5 videos a week: that's the physical limit. With TTS it's 20-30 videos of the same quality in a single overnight run. The voice is always in the same mood, never gets tired and never asks for a retake.
Scenario 2: Localization
Translating a video into 10 languages with live narrators means 10 narrators and 10 schedules. With Eleven v4, the same voice speaks dozens of languages (90+ according to ElevenLabs as of October 2026), and the pipeline itself takes hours. A native speaker checks pronunciation and translation before publishing.
Scenario 3: Audiobooks and podcasts
An 80,000-word book takes a narrator 8-10 hours of recording plus editing. Synthesis uses up your plan's credits: count the characters in the manuscript and compare that with your plan's limit (current prices: What's current). Claude reads the manuscript, adapts it for listening, and ElevenLabs voices it chapter by chapter.
ElevenLabs: how it works and how it's priced
ElevenLabs is one of the best-known professional TTS platforms.
What's inside (as of October 2026):
| Parameter | Value |
|---|---|
| Models | Eleven v4 (released September 28, 2026), previously Multilingual v2 was the main one; there are fast Flash and Turbo models |
| Languages | 90+ for Eleven v4 (according to ElevenLabs) |
| Output formats | several, set with the output_format parameter (for example, mp3_44100_128) |
| Cloning | instant clone and professional clone |
| Beyond voiceover | speech recognition, sound effects, music, voice agents |
Plans as of October 2026 (current prices: What's current and the ElevenLabs pricing page):
| Plan | Price per month | Credits per month | Commercial license | Cloning |
|---|---|---|---|---|
| Free | $0 | 10,000 | no | no |
| Starter | $6 | see the pricing page | yes | instant clone |
| Creator | $22 (first month $11) | see the pricing page | yes | instant and professional clone |
| Pro | $99 | see the pricing page | yes | instant and professional clone |
Credit usage depends on the model and the length of the text: measure it on a short test and multiply by your volume. The free plan has no commercial license, so for client projects you need a paid one.
Voice cloning: a photograph of your voice
Voice cloning is creating a digital clone of a voice from an audio sample. Once it's cloned, the system synthesizes speech with the same timbre, intonation and rhythm.
Sample requirements:
- Instant clone: a short sample of clean speech (ElevenLabs says Eleven v4 can clone from 10 seconds of recording). The cleaner and more varied the recording, the better the result
- Professional clone: a much longer recording (see the ElevenLabs help center for exact requirements)
- Format: WAV or MP3, no background noise
- No music, no other voices, no echo
Instant cloning, step by step:
- Go to elevenlabs.io → the Voices section → add a voice → instant clone (requires Starter or higher)
- Upload the audio file
- Give the voice a name (for example "MyVoice_EN")
- Confirm that the voice is yours or that you have its owner's permission
- Wait for processing and copy the Voice ID from the voice settings
After cloning, the voice is available through the API. You need the Voice ID for every request.
Important: ElevenLabs requires the voice owner's consent for cloning. Don't clone voices without permission: it violates the ToS (Terms of Service) and the laws of many countries.
A basic example: Claude writes the script, ElevenLabs voices it
from elevenlabs.client import ElevenLabs
import anthropic
import os
anthropic_client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
# Step 1: generate the script with Claude
response = anthropic_client.messages.create(
model="claude-sonnet-5-5", # current model IDs: see the Anthropic documentation
max_tokens=1000,
messages=[{
"role": "user",
"content": (
"Write a 60-second script for an ad for AI consulting services. "
"Tone: professional, confident, no empty promises. "
"About 150-160 words. Only the text to be read aloud, no stage directions."
)
}]
)
script = "".join(b.text for b in response.content if b.type == "text")
print(f"Script ({len(script.split())} words):\n{script}\n")
# Step 2: voice it with ElevenLabs
audio = eleven.text_to_speech.convert(
text=script,
voice_id="JBFqnCBsd6RMkjVDRZzb", # a voice from the docs example; get your own Voice ID in the Voices section
model_id="eleven_v4", # model as of October 2026; check the documentation for current model names
output_format="mp3_44100_128",
)
with open("ad-script.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
print("File saved: ad-script.mp3")Installation:
pip install elevenlabs anthropicEnvironment variables:
export ANTHROPIC_API_KEY="sk-ant-..."
export ELEVENLABS_API_KEY="sk_..."ElevenLabs MCP: voiceover straight from Claude Code
ElevenLabs offers official MCP servers. As of October 2026, the recommended option is the hosted server: nothing to install on your computer, and you sign in with OAuth. The older local server in the ElevenLabs repository is marked as deprecated. Once connected, Claude can voice text, pick voices and save files right from the chat, with no Python scripts (program code).
Connecting:
# ElevenLabs hosted server (recommended; OAuth sign-in, no API key needed)
claude mcp add --transport http elevenlabs https://api.elevenlabs.io/v1/mcp
# then inside Claude Code: /mcp and complete the sign-in
# The older local server (marked deprecated in the repository; requires uv and an API key in the environment)
claude mcp add elevenlabs --env ELEVENLABS_API_KEY=your_key -- uvx elevenlabs-mcpOr through .mcp.json (local server):
{
"mcpServers": {
"elevenlabs": {
"command": "uvx",
"args": ["elevenlabs-mcp"],
"env": {
"ELEVENLABS_API_KEY": "${ELEVENLABS_API_KEY}"
}
}
}
}Once connected, Claude gets tools (the set depends on the server; see elevenlabs.io/mcp and the server's repository for the exact list):
- Text-to-speech synthesis
- Voice cloning and voice management
- Transcription, audio cleanup, speech-to-speech conversion
- Generating soundscapes and music
An example request in the chat:
Voice the following text with one of the available voices and save it to intro.mp3: "Welcome to the lesson on voiceover. Over the next ten minutes you'll build your first pipeline: a script, a voice and a finished audio file."
Claude will call the MCP tool and save the file with no extra code.
Free alternatives: when ElevenLabs is overkill
| Service | Quality | Price | Cloning | Offline |
|---|---|---|---|---|
| ElevenLabs | ⭐⭐⭐⭐⭐ | Free and paid plans (see above) | ✅ on paid plans | ❌ |
| OpenAI TTS | ⭐⭐⭐⭐ | pay by volume through the API (prices: What's current) | ❌ | ❌ |
| Kokoro TTS | ⭐⭐⭐ | Free | ❌ (preset voices) | ✅ |
| macOS say | ⭐⭐ | Free | ❌ | ✅ |
OpenAI TTS: the API works like other OpenAI endpoints (API access points), comes with 13 built-in voices, and handles English and many other languages (according to OpenAI's documentation). No cloning of your own voice. A good fit when you already use OpenAI.
import openai, os
client = openai.OpenAI(api_key=os.environ["OPENAI_API_KEY"])
with client.audio.speech.with_streaming_response.create(
model="gpt-4o-mini-tts", # tts-1 and tts-1-hd are earlier models
voice="coral", # one of the built-in voices, see the list in OpenAI's documentation
input="Your voiceover text goes here",
instructions="Speak calmly and in a friendly way.",
) as response:
response.stream_to_file("output.mp3")Kokoro TTS: open source, runs locally, weights under the Apache 2.0 license. The quality is lower than ElevenLabs, but it's free and sends no data to the cloud. Note: according to the project README, Kokoro supports American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Chinese. For other languages, Kokoro won't work.
pip install "kokoro>=0.9.4" soundfilefrom kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code="a") # "a" is American English
generator = pipeline("Text for local voiceover", voice="af_heart")
for i, (_, _, audio) in enumerate(generator):
sf.write(f"output-{i}.wav", audio, 24000)macOS say: built into every Mac, free, works offline. The quality is robotic, but for prototypes and testing it's perfect.
# Basic voiceover
say "Hi, this is a test voiceover"
# Save to a file
say -v "Samantha" -o output.aiff "Text to read aloud"
# Convert to MP3 with ffmpeg
say -v "Samantha" -o /tmp/output.aiff "Text" && \
ffmpeg -i /tmp/output.aiff output.mp3 -y -loglevel error
# List the available US English voices
say -v "?" | grep en_USA YouTube pipeline with TTS voiceover
The full script: a video topic goes in, a finished MP4 (video file) comes out.
#!/usr/bin/env bash
# tts-video-pipeline.sh
# Usage: ./tts-video-pipeline.sh "Video topic" background.jpg
set -euo pipefail
TOPIC="$1"
BACKGROUND="${2:-background.jpg}"
OUTPUT_DIR="./output"
mkdir -p "$OUTPUT_DIR"
echo "Topic: $TOPIC"
# Step 1: Claude writes the script (via the claude CLI)
echo "Writing the script..."
SCRIPT=$(claude -p "Write a 3-minute YouTube script on: $TOPIC.
Tone: conversational, specific, no clichés.
Length: 400-450 words. Only the narrator's text, no stage directions or headings.")
echo "$SCRIPT" > "$OUTPUT_DIR/script.txt"
echo "Script: $(echo "$SCRIPT" | wc -w) words"
# Step 2: ElevenLabs voices it
echo "Recording the voiceover..."
python3 << PYEOF
from elevenlabs.client import ElevenLabs
import os
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
with open("$OUTPUT_DIR/script.txt", encoding="utf-8") as f:
text = f.read()
audio = client.text_to_speech.convert(
text=text,
voice_id="JBFqnCBsd6RMkjVDRZzb", # a voice from the docs example, replace it with your Voice ID
model_id="eleven_v4",
output_format="mp3_44100_128",
)
with open("$OUTPUT_DIR/voiceover.mp3", "wb") as out:
for chunk in audio:
out.write(chunk)
print("Voiceover ready")
PYEOF
# Step 3: FFmpeg builds the video
echo "Assembling the video..."
DURATION=$(ffprobe -v error -show_entries format=duration \
-of csv=p=0 "$OUTPUT_DIR/voiceover.mp3" | awk '{print int($1+1)}')
ffmpeg \
-loop 1 -i "$BACKGROUND" \
-i "$OUTPUT_DIR/voiceover.mp3" \
-c:v libx264 -tune stillimage \
-c:a aac -b:a 192k \
-pix_fmt yuv420p \
-t "$DURATION" \
"$OUTPUT_DIR/video.mp4" \
-y -loglevel error
echo "Done: $OUTPUT_DIR/video.mp4 (${DURATION}s)"Running it:
chmod +x tts-video-pipeline.sh
./tts-video-pipeline.sh "How RAG works in Claude" background.jpgA voice journal and a podcast in 5 minutes
You speak your thoughts out loud → Whisper (STT, Speech-to-Text, turning speech into text) transcribes them → Claude edits them for listening → ElevenLabs gives them a professional voiceover.
import whisper
import anthropic
from elevenlabs.client import ElevenLabs
import os
# Models
whisper_model = whisper.load_model("small")
claude = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
def voice_to_podcast(audio_file: str, output_file: str = "podcast.mp3"):
"""
audio_file: a voice recording (WAV/MP3)
output_file: the finished podcast episode
"""
# Step 1: transcribe the voice note
print("Transcribing...")
result = whisper_model.transcribe(audio_file, language="en")
raw_text = result["text"]
print(f"Transcript ({len(raw_text.split())} words): {raw_text[:100]}...")
# Step 2: Claude edits it for listening
print("Editing...")
response = claude.messages.create(
model="claude-sonnet-5-5", # current model IDs: see the Anthropic documentation
max_tokens=2000,
messages=[{
"role": "user",
"content": (
"This is a transcript of spoken speech. Edit it for a podcast:\n"
"- Remove filler words (um, uh, like, you know)\n"
"- Fix broken-off sentences\n"
"- Keep the conversational tone and the structure of the thought\n"
"- Don't add anything new, only edit\n\n"
f"Text:\n{raw_text}"
)
}]
)
edited_text = "".join(b.text for b in response.content if b.type == "text")
# Step 3: ElevenLabs gives it a professional voiceover
print("Recording the voiceover...")
audio = eleven.text_to_speech.convert(
text=edited_text,
voice_id="JBFqnCBsd6RMkjVDRZzb", # a voice from the docs example, replace it with your Voice ID
model_id="eleven_v4",
output_format="mp3_44100_128",
)
with open(output_file, "wb") as out:
for chunk in audio:
out.write(chunk)
print(f"Podcast ready: {output_file}")
return edited_text
# Run it
edited = voice_to_podcast("my-thoughts.wav", "podcast-ep01.mp3")
print("\nEdited text saved to podcast-ep01-script.txt")
with open("podcast-ep01-script.txt", "w", encoding="utf-8") as f:
f.write(edited)Services built on TTS
TTS isn't just a tool for yourself. It's the foundation for three kinds of services. There are no prices here on purpose: work out the price of your service with the lesson How to set a price, and nobody can guarantee income.
Option 1: Voicing books and courses
The client brings a manuscript, you deliver audio chapter by chapter. Claude edits the text for listening, ElevenLabs voices it. Your cost is calculated in credits: the number of characters in the manuscript against your plan's credits. You need a paid plan with a commercial license and confirmation that the client holds the rights to the text.
The idea: dozens of hours of a narrator's work are replaced by a few hours of your pipeline. Check the quality by ear: mistakes in stress and in proper names are still your responsibility.
Option 2: Content localization
A creator or a business wants to translate a video course into several languages. You use Eleven v4 with the same voice. A native speaker checks translation and pronunciation, otherwise the mistakes will go out with the publication.
Option 3: TTS SaaS (Software-as-a-Service) for small businesses
Realtors, tutors, small businesses all record announcements and presentations by voice. You can wrap the ElevenLabs API in a simple web interface: upload text → pick a voice → download the MP3. Before launching, check the ElevenLabs API terms on resale and commercial use, and work out your costs with the lesson What a client costs and what they bring in.
Practice
Sign up at elevenlabs.io → get an API key (the Free plan is enough to start)
Install the dependencies:
bash pip install elevenlabs anthropicRun the basic "Claude writes → ElevenLabs voices" script from the theory section. Listen to the result.
Clone your voice (requires Starter or higher): record 60 seconds of clear speech (read any text), upload it to ElevenLabs in the Voices section (instant clone). Replace
voice_idin the code with your Voice ID.Create the
tts-video-pipeline.shfile, make it executable, run it with any topic. Check the resulting MP4.(Advanced) Install Kokoro TTS locally, voice an English text, and compare it with ElevenLabs. Note the difference in quality.
Tools and resources
- ElevenLabs: sign-up, API key, voice cloning
- ElevenLabs API Docs: full documentation
- ElevenLabs MCP: the hosted MCP server; the older local one: elevenlabs-mcp
- OpenAI TTS: an alternative with 13 built-in voices
- Kokoro TTS: open source, offline, a limited set of languages
- OpenAI Whisper: for the STT part of the pipeline
- Prices and versions: What's current
Key takeaways
TTS is the last manual step in the content pipeline. ElevenLabs closes it. One voice → many videos, 90+ languages with Eleven v4, no studio.
Voice cloning = a short sample → endless voiceover. Record once, and your voice works without you, as long as you have the voice owner's permission.
Price TTS services (book narration, localization, TTS SaaS) from your costs: plan credits, your time, a native speaker's review. Nobody guarantees prices or income.
A free starter stack: macOS say for tests, Kokoro TTS for offline prototypes in English, ElevenLabs Free (10,000 credits a month, no commercial license) for your first experiments.
Next lesson
→ Music with AI: Suno and sound design: background music, jingles and sound effects for videos and podcasts
The mark stays in this browser only and is never sent anywhere. My progress