elevenlabs-core-workflow-b
Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text. Use when converting one voice to another, generating sound effects from a text description, removing background noise from a recording, or transcribing audio. Trigger with "elevenlabs speech to speech", "voice changer", "sound effects", "audio isolation", "remove background noise", "elevenlabs transcribe".
Allowed Tools
Provided by Plugin
elevenlabs-pack
Claude Code skill pack for ElevenLabs (18 skills)
Installation
This skill is included in the elevenlabs-pack plugin:
/plugin install elevenlabs-pack@claude-code-plugins-plus
Click to copy
Instructions
ElevenLabs Core Workflow B — Speech-to-Speech, Sound Effects & Audio Isolation
Overview
Secondary ElevenLabs workflows beyond TTS: (1) Speech-to-Speech voice conversion,
(2) Sound Effects generation from text descriptions, (3) Audio Isolation for noise
removal, and (4) Speech-to-Text transcription. Each maps to one API endpoint and
has both a TypeScript SDK and a cURL path.
Full code for every step lives in references/implementation.md;
copy-ready invocations are in references/examples.md.
Prerequisites
- Completed
elevenlabs-install-authsetup. - For STS: source audio file in MP3/WAV/M4A format.
- For audio isolation: noisy audio file to clean.
Authentication
The SDK client (new ElevenLabsClient()) reads the API key from the
ELEVENLABSAPIKEY environment variable automatically — never hardcode it. cURL
requests send it as the xi-api-key: ${ELEVENLABSAPIKEY} header. Full auth setup
is covered by the elevenlabs-install-auth skill.
Instructions
Import the SDK once, then call the relevant module. The client authenticates from
the environment:
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";
import { Readable } from "stream";
import { pipeline } from "stream/promises";
const client = new ElevenLabsClient();
- Speech-to-Speech (voice changer) —
client.speechToSpeech.convert(voiceId, …)
against POST /v1/speech-to-speech/{voiceid}. Use modelid: "elevenenglishsts_v2"
and set removebackgroundnoise: true for built-in cleanup.
- Sound Effects —
client.textToSoundEffects.convert({ text, … })against
POST /v1/sound-generation. Tune duration_seconds (0.5–30) and
prompt_influence (0–1; higher follows the prompt more closely).
- Audio Isolation —
client.audioIsolation.audioIsolation({ audio })against
POST /v1/audio-isolation, or the streaming variant for large files.
- Speech-to-Text —
client.speechToText.convert({ audio, modelid: "scribev1" })
against POST /v1/speech-to-text; optionally enable diarize and word timestamps.
Each returns an audio stream (steps 1–3) piped to disk, or a transcript object
(step 4). See references/implementation.md for the
complete helper functions and cURL equivalents.
First example — Speech-to-Speech skeleton
async function speechToSpeech(sourceAudioPath, targetVoiceId, outputPath) {
const audio = await client.speechToSpeech.convert(targetVoiceId, {
audio: createReadStream(sourceAudioPath),
model_id: "eleven_english_sts_v2",
voice_settings: JSON.stringify({ stability: 0.5, similarity_boost: 0.8 }),
remove_background_noise: true,
});
await pipeline(Readable.fromWeb(audio as any), createWriteStream(outputPath));
}
API Endpoint Summary
| Feature | Method | Endpoint | Billing |
|---|---|---|---|
| Speech-to-Speech | POST | /v1/speech-to-speech/{voice_id} |
Per character |
| Sound Effects | POST | /v1/sound-generation |
Per generation |
| Audio Isolation | POST | /v1/audio-isolation |
1,000 chars/min of audio |
| Audio Isolation Stream | POST | /v1/audio-isolation/stream |
1,000 chars/min of audio |
| Speech-to-Text | POST | /v1/speech-to-text |
Per audio minute |
Output
- Steps 1–3 write an audio file to the
outputPathyou pass and log a
confirmation, e.g. Voice-converted audio saved to converted.mp3 or
Clean audio saved to clean_interview.mp3.
- Step 4 returns a transcript object:
result.textholds the full
transcription, and result.words (when present) carries word-level
{ start, end, text } timestamps.
- cURL paths stream the resulting audio directly to the
--outputfile.
Error Handling
| Error | HTTP | Cause | Solution |
|---|---|---|---|
modelcannotdovoice_conversion |
400 | Wrong model for STS | Use elevenenglishsts_v2 |
audiotooshort |
400 | STS input under 1 second | Use longer audio clip |
audiotoolong |
400 | STS input over limit | Trim to under 5 minutes |
invalidsoundprompt |
400 | Nonsensical SFX description | Write descriptive, specific prompts |
filetoolarge |
413 | Audio isolation over 500MB | Compress or split the file |
quota_exceeded |
401 | Character/generation limit hit | Check usage dashboard |
Examples
Worked, copy-ready invocations for all four workflows — including the three
sound-effect variants (rain, laser, seamless forest loop), the "Rachel" voice
conversion, an audio-isolation clean-up, and a transcription with word timestamps —
are in references/examples.md. A one-liner:
// Generate a 10-second rain sound effect, faithful to the prompt
await generateSoundEffect(
"Heavy rain on a tin roof with distant thunder",
"rain.mp3",
{ duration: 10, promptInfluence: 0.6 }
);
Resources
- Full implementation walkthrough — every step's
SDK + cURL code, sound-effect tips, and audio-isolation limits.
- Worked examples — copy-ready invocations per workflow.
- Speech-to-Speech API
- Sound Effects API
- Audio Isolation API
- Speech-to-Text API
Next Steps
For common errors, see elevenlabs-common-errors. For SDK patterns, see elevenlabs-sdk-patterns.