D

Deepgram

GitHub 资料 · @deepgram

Build a real-time voice agent on Deepgram's Voice Agent API: one WebSocket at wss://agent.deepgram.com/v1/agent/converse that runs speech-to-text, a language model (Deepgram-hosted or your own endpoint), and text-to-speech, with barge-in, mid-call updates, and function calling that your own client executes. Use when someone says "voice agent", "voice bot", "speech-to-speech", "talk to an AI on the phone", "agent.deepgram.com", "Settings message", "FunctionCallRequest", "barge-in", "Twilio voice agent", or asks whether to build on Deepgram directly or through LiveKit Agents, Pipecat, Vapi, or Retell. Routes to the api, docs, starters, recipes, examples, and per-language SDK skills for the full reference.
deepgram/voice-agent
Turn text into spoken audio with Deepgram. Use when someone asks for text-to-speech, TTS, speech synthesis, a synthetic voice, a voice for a voice agent, an IVR prompt, or an audio version of some text, and whenever they mention Aura, Aura-2, Flux TTS, /v1/speak or /v2/speak. Trigger phrases: "text to speech", "TTS", "speak endpoint", "generate speech", "synthesize audio", "read this aloud", "which Deepgram voice", "Aura vs Flux TTS", "TTS with barge-in". Gets an agent to a correct first request, then routes to the api, docs, starters, recipes, examples, setup-mcp and per-language SDK text-to-speech skills.
deepgram/text-to-speech
Analyze text you already have with Deepgram's Read API. Use when a task says "text intelligence", "read API", "/v1/read", "analyze text", "sentiment of this text", "summarize this transcript", "summarize a document", "topic detection on text", "intent recognition on text", "analyze a support ticket", or "analyze a chat log". One REST call, POST /v1/read, with four features: summarize, sentiment, topics, intents. Covers the two required query parameters people miss, the text-versus-url body, the English-only limit, and why entity detection is not here. Routes to audio-intelligence for audio input and to the api, docs, recipes, starters, and per-language SDK skills.
deepgram/text-intelligence
Replace with description of the skill and when to use it.
deepgram/template-skill
Clone a ready-to-run Deepgram demo app and start building on top of it. Use whenever someone wants a quick working demo, needs to prototype with Deepgram, or is starting a new project that uses speech-to-text, text-to-speech, voice agents, audio intelligence, or live streaming. Match the user's language, framework, and desired Deepgram feature to the right starter.
deepgram/starters
Start here for Deepgram speech-to-text. Use when a task says "speech to text", "STT", "transcribe", "transcription", "live transcription", "captions", "diarization", "nova-3", "Flux", "turn detection", or "end of turn". Picks the model family (Nova on /v1/listen for general transcription, Flux STT on /v2/listen for conversational audio with built-in turn detection), gets a first request working, and routes to the api, docs, recipes, starters, examples, and per-language SDK skills for everything else.
deepgram/speech-to-text
Set up a Deepgram MCP server for your AI coding tool. Offers three paths: the Deepgram CLI MCP proxy (dg mcp), the standalone deepgram-mcp package, and the credential-free hosted documentation MCP. Use whenever someone wants to install Deepgram's agentic tools, set up the MCP server, or connect their editor to Deepgram.
deepgram/setup-mcp
Run Deepgram in infrastructure you control. Use when a task says "self-hosted", "self hosting", "on-prem", "on-premise", "air-gapped", "airgapped", "private deployment", "BYOC", "docker compose deepgram", "podman", "kubernetes", "helm chart", "sagemaker", "license proxy", "quay.io", "distribution credentials", "data residency", or "FIPS". Decides whether self-hosting is even the right answer versus a regional endpoint or Deepgram Dedicated, walks the licensing and container credential bootstrap, and routes to per-target guidance for Docker/Podman, Kubernetes/Helm, and Amazon SageMaker.
deepgram/self-hosted
Find focused, runnable Deepgram recipes for a specific feature × language. Use whenever someone wants a minimal working code snippet for ONE feature (transcribe URL, diarize, smart-format, voice agent connect, etc.) rather than a full starter app. Recipes are under 50 lines, read DEEPGRAM_API_KEY from env, and ship with a runnable example_test. Covers Python, JavaScript, Go, .NET, Java, Rust, and the Deepgram CLI.
deepgram/recipes
Find working Deepgram integration examples with third-party platforms and frameworks. Use whenever someone wants to integrate Deepgram with Twilio, LiveKit, LangChain, Vercel AI SDK, Discord, Vonage, Pipecat, Expo, FastAPI, Cloudflare Workers, Slack, Telegram, LlamaIndex, Zoom, Next.js, Nuxt, Django, SvelteKit, NestJS, Spring Boot, CrewAI, Riverside, SignalWire, and more. Examples are full runnable integration demos, not minimal feature snippets.
deepgram/examples
Find the right Deepgram documentation for any task. Use whenever someone needs help locating docs, understanding which API to use, or wants to ask questions about Deepgram. Covers all product areas: speech-to-text (Nova, Flux STT), text-to-speech (Aura, Flux TTS), voice agents, audio intelligence, and self-hosted deployments.
deepgram/docs
Drive Deepgram from the terminal with the official CLI. Use when a task says "deepgram cli", "dg command", "deepctl", "dg listen", "dg speak", "transcribe from terminal", "transcribe a file from the command line", "caption a file with the CLI", "dg init", "scaffold a starter", "dg mcp", "dg skills", "dg login", "dg keys", "dg usage", or asks to script Deepgram in CI or a shell pipeline. Covers install, authentication, the real command surface, and the gaps where you should reach for the API or an SDK instead.
deepgram/cli
Run a Deepgram voice agent in the browser with the four Browser Agent SDK packages: @deepgram/agents (core WebSocket session, mic, player), @deepgram/react (AgentProvider and hooks), @deepgram/ui (pre-built React components), and @deepgram/agents-widget (drop-in, no framework). Use when a task says "browser voice agent", "voice widget", "embed a voice agent", "react voice agent", "@deepgram/react", "@deepgram/ui", "agents-widget", "AgentProvider", "useAgentState", "useDeepgramAgent", "AgentSession", "tokenFactory", "Orb", "voice agent on my website", or asks how to keep a Deepgram API key out of client-side code. Picks the layer, gets one path running, and routes to the voice-agent skill for the WebSocket contract underneath.
deepgram/browser-agent
Analyze what was said in audio, not just transcribe it. Use when a task mentions "audio intelligence", "sentiment", "sentiment analysis", "summarize a recording", "summarization", "topics", "topic detection", "intents", "intent recognition", "entity detection", "detect entities", "extract names and amounts from a call", or "analyze a call recording". These are five query parameters layered on the speech-to-text endpoint /v1/listen (summarize, sentiment, topics, intents, detect_entities), not a separate API. Covers the prerecorded-only and English-only limits, where each result lives in the JSON, and the errors you get when you cross a limit. Routes to the api, docs, recipes, starters, and per-language SDK skills.
deepgram/audio-intelligence
Deepgram API reference for speech-to-text, text-to-speech, voice agents, audio intelligence, and account management. Use whenever building with Deepgram APIs — REST or WebSocket. Covers authentication, all endpoints, query parameters, request/response schemas, and WebSocket message formats. Reference files are organized by domain: listen (STT — Nova and Flux STT), speak (TTS — Aura and Flux TTS), agent (voice agents), read (text/audio intelligence), models, projects, auth, and self-hosted.
deepgram/api