inworld

v2026.09.24

Inworld TTS API. Covers voice cloning, audio markups, timestamps. Use when integrating Inworld text-to-speech, cloning voices, adding audio markups (SSML-like), or aligning viseme timestamps. Keywords: Inworld, text-to-speech, TTS, voice cloning, visemes.

GitHub
Install command
npx skhub add itechmeat/inworld
Markdown
SKILL.md

Inworld AI

Text-to-Speech platform with voice cloning, audio markups, and timestamp alignment.

Quick Navigation

TopicReference
Installationinstallation.md
Voice Cloningcloning.md
Voice Controlvoice-control.md
API Referenceapi.md

When to Use

  • Text-to-speech audio generation
  • Voice cloning from 5-15 seconds of audio
  • Emotion-controlled speech ([happy], [sad], etc.)
  • Word/phoneme timestamps for lip sync
  • Custom pronunciation with IPA

Models

ModelIDLatencyPrice
TTS-2 Flashinworld-tts-2-flashlowestsee pricing
TTS-2inworld-tts-2latestsee pricing
TTS 1.5 Maxinworld-tts-1.5-maxlegacylegacy
TTS 1.5 Miniinworld-tts-1.5-minilegacylegacy

Minimal Example

import requests, base64, os

response = requests.post(
    "https://api.inworld.ai/tts/v1/voice",
    headers={"Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}"},
    json={"text": "Hello!", "voiceId": "Ashley", "modelId": "inworld-tts-1.5-max"}
)
audio = base64.b64decode(response.json()['audioContent'])

Key Features

  • 15 languages — en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar
  • Instant cloning — 5-15 seconds audio, no training
  • Audio markups — [happy], [laughing], [sigh] (English only)
  • Timestamps — word, phoneme, viseme timing for lip sync
  • Streaming — /voice:stream endpoint
  • TTS-2 steering — natural-language bracketed directions such as [say excitedly] or [whisper in a hushed style]
  • Delivery mode — STABLE, BALANCED, CREATIVE trade consistency for emotional range
  • Cross-lingual synthesis — reuse one voice across multiple languages; voice localization improves native-sounding output

Release Highlights (TTS-2)

  • Realtime TTS-2 becomes the new primary model line via modelId="inworld-tts-2".
  • Steering moves beyond the older fixed emotion tags: free-form bracketed directions can control style, pitch, speed, intensity, and non-verbals.
  • Multilingual coverage expands with production quality across 15 languages and broader experimental coverage beyond that.
  • deliveryMode adds a stability-vs-creativity knob, and specifying language matters more for cross-lingual output quality.

Release Updates (August 2026)

  • New inworld-tts-2-flash model: lowest latency and cost, full language coverage, instant voice cloning, and timestamp alignment. Steering and Professional Voice Cloning remain exclusive to inworld-tts-2.
  • Steering instructions now persist until explicitly changed: a reserved [reset] tag ends a styled passage, and a <break/> pause no longer clears the active instruction. A request-level instruction field on Synthesize Speech is an alternative to inline tags.
  • Cross-lingual voice synthesis and voice localization improve native-sounding output when one voice is reused across languages.

Prohibitions

  • Audio markups work only in English
  • Use ONE emotion markup at text beginning
  • Match voice language to text language
  • Instant cloning may not work for children's voices or unique accents

Links

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/inworld

Default branch

master

Latest commit

7ae8a00

Tree SHA

47f5439