elevenlabs-core-workflow-a
Implement ElevenLabs text-to-speech and voice cloning workflows. Use when building TTS features, cloning voices from audio samples, streaming speech to a chatbot, or implementing the primary ElevenLabs money-path: voice generation. Trigger with "elevenlabs TTS", "text to speech", "voice cloning elevenlabs", "clone a voice", "generate speech", "elevenlabs voice".
Allowed Tools
Provided by Plugin
elevenlabs-pack
Claude Code skill pack for ElevenLabs (18 skills)
Installation
This skill is included in the elevenlabs-pack plugin:
/plugin install elevenlabs-pack@claude-code-plugins-plus
Click to copy
Instructions
ElevenLabs Core Workflow A — TTS & Voice Cloning
Overview
The primary ElevenLabs workflows: (1) Text-to-Speech with voice settings, (2) Instant Voice Cloning from audio samples, (3) streaming TTS via WebSocket for real-time applications, and (4) voice-library management. This SKILL.md walks the full flow at a high level and carries the first TTS example inline; the deep code for cloning, streaming, and management lives in the full implementation walkthrough.
Prerequisites
- Completed
elevenlabs-install-authsetup - Valid API key with sufficient character quota
- For voice cloning: audio recording(s) of the target voice (min 30 seconds, clean audio)
Instructions
Step 1: Advanced Text-to-Speech
Instantiate the client, call textToSpeech.convert(voiceId, opts), and pipe the returned stream to a file. The voice_settings block is where you tune delivery:
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
import { Readable } from "stream";
import { pipeline } from "stream/promises";
const client = new ElevenLabsClient();
async function generateSpeech(
text: string,
voiceId: string,
outputPath: string
) {
const audio = await client.textToSpeech.convert(voiceId, {
text,
model_id: "eleven_multilingual_v2",
voice_settings: {
stability: 0.5, // Lower = more expressive, higher = more consistent
similarity_boost: 0.75, // How closely to match the original voice
style: 0.3, // Amplify the speaker's style (adds latency if > 0)
speed: 1.0, // 0.7 to 1.2 range
},
// Optional: enforce language for multilingual model
// language_code: "en", // ISO 639-1
});
await pipeline(Readable.fromWeb(audio as any), createWriteStream(outputPath));
console.log(`Generated: ${outputPath}`);
}
await generateSpeech("Welcome to our platform.", "21m00Tcm4TlvDq8ikWAM", "stable.mp3");
Step 2: Instant Voice Cloning (IVC)
Clone a voice from 1-25 audio samples with client.voices.add({ name, description, files }), which returns a voiceid you can use immediately in textToSpeech.convert. Use similarityboost: 0.85 on cloned voices to stay close to the original. Full cloneVoice implementation: implementation.md, Step 2.
Step 3: WebSocket Streaming TTS
For real-time apps (chatbots, live narration), open wss://api.elevenlabs.io/v1/text-to-speech/{voiceId}/stream-input with the low-latency elevenflashv2_5 model. Send a space as Beginning-of-Stream, stream text chunks, then an empty string as End-of-Stream; collect base64 audio frames until isFinal. Full streamTTSWebSocket implementation: implementation.md, Step 3.
Step 4: Voice Management
List, inspect, update, and delete voices with client.voices.getAll(), getSettings, editSettings, and delete. Full helpers: implementation.md, Step 4.
Tuning Reference
Two lookup tables — the voice-cloning input requirements and the full
voice_settings range/effect guide with per-use-case starting points — live in
implementation.md. Quick defaults:
- Narration:
stability=0.5, similarity_boost=0.75, style=0.0 - Conversational:
stability=0.4, similarity_boost=0.6, style=0.3 - Cloned voice:
stability=0.5, similarity_boost=0.85, style=0.0
Output
- Text-to-Speech (Step 1): an audio stream written to
outputPath(e.g.stable.mp3); console logsGenerated:plus the output path. - Voice cloning (Step 2): a new
voice_id(logged asCloned voice created:plus the id) plus an immediately-usable audio stream in the cloned timbre. - WebSocket streaming (Step 3): a concatenated
Bufferof base64-decoded audio chunks assembled as frames arrive. - Voice management (Step 4): printed voice listings (name, voice_id, category), current/updated settings, or a delete confirmation.
Error Handling
| Error | HTTP | Cause | Solution |
|---|---|---|---|
voicenotfound |
404 | Invalid voice_id | List voices first: GET /v1/voices |
texttoolong |
400 | Over 5,000 chars per request | Split text and use previoustext/nexttext for prosody |
quota_exceeded |
401 | Character limit reached | Check usage, upgrade plan |
toomanyconcurrent_requests |
429 | Exceeds plan concurrency | Queue requests; see concurrency limits |
invalidvoicesample |
400 | Bad audio file for cloning | Use clean audio, supported format, 30s+ |
WebSocket modelnotsupported |
N/A | eleven_v3 not available for WS | Use elevenflashv25 or elevenmultilingual_v2 |
Examples
Four complete input-to-audio scenarios are in references/examples.md:
- Generate narration from a script — batch a marketing script to one MP3 with a premade voice.
- Clone a narrator voice and speak with it — clone from two samples, then synthesize with the returned
voice_id. - Stream an LLM response as speech — pipe chatbot chunks through the WebSocket for real-time playback.
- Audit and prune your voice library — list every voice by category, then delete a stale clone.
Resources
Next Steps
For speech-to-speech, sound effects, and audio isolation, see the companion skill elevenlabs-core-workflow-b, which covers the remaining ElevenLabs audio-transformation endpoints.