nemotron-voice-agent-builder

v2026.09.24

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Use when building, scaffolding, or iterating on a real-time voice-agent pipeline, including speech (ASR/TTS) customization and cloud or local deployment. Not for offline/batch speech-to-text, text-only chat or RAG, or generic Docker or CUDA work unrelated to a voice agent.

GitHub
Install command
npx skhub add nvidia/nemotron-voice-agent-builder
Markdown
SKILL.md

Create Voice Agent

Creates a working NVIDIA voice agent or updates an existing project:

  • Cascaded: ASR transcribes, a text LLM answers, TTS speaks.
  • Omni: one audio-in LLM replaces ASR and the text LLM. TTS still speaks.

When to Use This Skill

Use this skill when the user wants to build, scaffold, configure, refine, or fix an NVIDIA voice agent — a real-time speech pipeline with audio input and spoken output — on Pipecat or LiveKit. This covers Cascaded (ASR → LLM → TTS) and Omni (audio-in LLM + TTS) pipelines, speech customization (ASR word boosting, TTS pronunciation), multilingual routing, and cloud or local deployment (NIM, vLLM, NeMo-Speech.cpp) on workstations, DGX Spark, or Jetson Thor.

Do not use this skill for text-only chatbots or RAG, standalone batch speech-to-text transcription, generic Docker or infrastructure help, or unrelated CUDA or model work that has no voice-agent pipeline.

Workflow

Classify the starting state before following the phases:

  • For an empty project, follow all phases in order.
  • For a working existing project, read references/operations/iterate.md first, and then route only the references needed for the requested change.
  • For a broken existing project, read references/operations/troubleshoot.md first. After restoring the baseline, continue with references/operations/iterate.md.

For an empty project, follow these phases in order:

PhaseRead
1. Intakereferences/intake.md resolves pipeline and framework first, then routes preflight, models, language, and domain files
2. Buildreferences/platforms/readiness.md if self-hosted → exact profile or quantization variant in the routed model file → references/output-contract.md → routed deployment path → references/operations/observability.md → selected framework files
3. Hand overreferences/operations/run.md gates the client on the generated scripts/smoke.sh, then references/operations/troubleshoot.md if the spoken exchange fails
4. Iteratereferences/operations/iterate.md for changes to a working agent

Wait for approval before writing files.

references/preflight.md §4 probes the host and selects the deployment path. The routed platform file owns that path from there.

Also read references/networking/remote-webrtc.md when Pipecat WebRTC crosses hosts or networks.

Rules

  • Resolve every model id, image tag, profile, and serve flag from current NVIDIA documentation at build time. Never from memory.
  • Read every routed file before generating code, and keep the language and behavior locks approved in intake.
  • Handover requires a successful spoken exchange. A running process is not a working voice agent.
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Apache-2.0

Source path

skills/nemotron-voice-agent-builder

Default branch

main

Latest commit

ef46204

Tree SHA

94ca43b