tts-generation

v2026.09.24

AI text-to-speech generation using OpenAI TTS, ElevenLabs, and Google TTS backends. Converts text to audio files with voice selection, speed control, and format options.

GitHub
安装命令
npx skhub add oimiragieo/tts-generation
Markdown
SKILL.md

TTS Generation

Overview

Generate speech audio from text using AI backends.

  • OpenAI TTS — tts-1 (low latency) / tts-1-hd (studio quality), 6 voices, 57 languages
  • ElevenLabs — eleven_turbo_v2 / eleven_multilingual_v2, cloneable voices, 29 languages
  • Google TTS — gTTS Python library, 40+ languages, free tier

Backend Comparison

FeatureOpenAI TTSElevenLabsGoogle TTS
QualityHighHighestMedium
LatencyLow (tts-1)MediumLow
Cost~$15/1M chars~$22/1M charsFree (limited)
Voices6 presetCloneable40+ languages
Max chars4096/requestUnlimited~5000/request
StreamingYesYesNo

Quick Start

OpenAI TTS (Recommended)

from pathlib import Path
from openai import OpenAI

client = OpenAI()

response = client.audio.speech.with_streaming_response.create(
    model="tts-1-hd",  # tts-1 for speed, tts-1-hd for quality
    voice="nova",       # alloy | echo | fable | onyx | nova | shimmer
    input="Hello world",
    speed=1.0,          # 0.25 to 4.0
)
response.stream_to_file(Path("output.mp3"))

ElevenLabs

from elevenlabs import ElevenLabs

client = ElevenLabs(api_key="YOUR_API_KEY")
audio = client.text_to_speech.convert(
    voice_id="21m00Tcm4TlvDq8ikWAM",  # Rachel
    model_id="eleven_turbo_v2",
    text="Hello world",
    output_format="mp3_44100_128",
)
with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

Google TTS (Free)

from gtts import gTTS
gTTS(text="Hello world", lang="en", slow=False).save("output.mp3")

Long-Text Chunking

For text exceeding limits, split at sentence boundaries and concatenate with pydub. Pattern: iterate sentences, accumulate into current until max_chars (4000), flush to chunks on overflow.

Output Formats

mp3 (general), opus (streaming), flac (lossless archival), wav (editing), pcm (raw pipeline).

Installation

pip install openai elevenlabs gtts pydub
export OPENAI_API_KEY="sk-..."
export ELEVENLABS_API_KEY="..."

Agent Usage Pattern

  • OpenAI TTS: documentation/demos narration
  • ElevenLabs: cloned voices or highest quality
  • Google TTS: multilingual free-tier
  • Chunk at sentence boundaries; cache by content hash

Related Skills

  • transcription — Reverse: audio to text via Whisper
  • ai-ml-expert — Advanced ML pipeline integration

Memory Protocol (MANDATORY)

Before starting: Read .claude/context/memory/learnings.md

After completing:

  • New pattern → .claude/context/memory/learnings.md
  • Issue found → .claude/context/memory/issues.md
  • Decision made → .claude/context/memory/decisions.md

ASSUME INTERRUPTION: If it's not in memory, it didn't happen.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

未指定

源路径

.claude/skills/tts-generation

默认分支

main

最新提交

64b580e

Tree SHA

42a1df4