Pydantic AI
Python agent framework for building production-grade GenAI applications with the "FastAPI feeling".
Quick Navigation
| Topic | Reference |
|---|---|
| Agents | agents.md |
| Capabilities | agents.md |
| Tools | tools.md |
| Models | models.md |
| Embeddings | embeddings.md |
| Evals | evals.md |
| Integrations | integrations.md |
| Graphs | graphs.md |
| UI Streams | ui.md |
| Installation | installation.md |
When to Use
- Building AI agents with structured output
- Need type-safe, IDE-friendly agent development
- Require dependency injection for tools
- Multi-model support (OpenAI, Anthropic, Gemini, etc.)
- Production observability with Logfire
- Complex workflows with graphs
Installation
See references/installation.md for full/slim install options and optional dependency groups. Requires Python 3.10+.
Release Highlights (2.40.0 -> 2.48.0)
- Security (upgrade) (
2.44.0, backported to1.107.6): four fixes reached throughweb_fetch_toolor OTel instrumentation — IPv6 zone-identifier bypass of the private-IP blocklist when local URLs are opted in, superlinearweb_fetchHTML/charset processing that could stall the event loop, domain blocklists compared without resolver normalization, and span content leaks withinclude_content=False(GHSA-vmxc-h2x2-jmf3, GHSA-fpf4-vwcp-v4hp, GHSA-22h6-qm39-v87j, GHSA-4x9p-g9wm-8q7f). - Realtime session control (
2.40.0/2.46.0):handle_barge_in=Truewithinterrupt(played_bytes=...),RealtimeSession.enqueue()for out-of-band prompts,respond=onsend(),wait_for_playback(), and@agent.on_eventlisteners. - Image generation (
2.41.0): directImageGeneratorAPI alongside the built-in tool;fallback_modelonImageGeneration/XSearchdeprecated in favor offallback_subagent_model. - Providers and models (
2.41.0/2.42.0/2.48.0):openai-codexprovider for ChatGPT/Codex subscription auth,GitHubCopilotProvider; OpenAIgpt-6-sol/gpt-6-lunaand Claude Opus 5.5 (claude-opus-5-5). - Pricing (
2.40.0):pydantic_ai.prices.update_in_background()refreshesgenai-pricesdata in the background. - Temporal (
2.46.0):event_stream_topiconTemporalDurabilitystreams agent events via Workflow Streams.
Release Highlights (2.23.0 -> 2.39.0)
- Realtime speech-to-speech (
2.28.0):Agent.realtime()with browser WebRTC plus server sideband support; Azure AI Voice Live via theazure_voice_livesetting (2.29.0);RealtimeSession.send_audio()accepts async iterables (2.36.0). - Run cancellation (
2.26.0):AgentRun.cancel(),RunContext.cancel(), andRunCancelled;stream_run_events()returns a publicAgentRunEventshandle withcancel()and run-state access. - Deferred tool reveal (
2.23.0/2.26.0/2.30.0):ToolAvailabilityDeltaPartwith nativetool_addition/additional_tools; tools can stay hidden until revealed via tool search,load_capability, orToolReturn.tools, on each provider's native deferral/addition channel; a deferred tool must be revealed before it can be called. - Cost tracking (
2.23.0):costonRunUsageandcost_limitonUsageLimits. Context window (2.38.0):context_windowonModelProfileandcontext_window_usedonRunContext. - Typed events (
2.38.0): application code and capabilities can emit typedCustomEvent/CapabilityEventinto the run event stream and subscribe with@on_event. Durable execution (2.36.0):@durable_operationfor capabilities plus a public backend API for third-party durable execution engines. - Providers and models: Crusoe (
2.28.0), Snowflake Cortex (2.27.0), and vLLM (2.38.0) providers;gemini-3.7-flash(2.30.0),gemini-3.8-flash(2.38.0), Claude Fable 5.1 / Mythos 5.1 (2.38.0), OpenAIgpt-6-astra(2.39.0), GLM-5.3 onZaiModel(2.34.0), DeepSeek V4 Flash (2.26.0). - MCP and HTTP clients:
MCPToolsetsupports FastMCP 4 and MCP SDK v2 alongside FastMCP 3 (2.29.0). Builds moved tohttpx2clients;pydantic-ai[anthropic]requiresanthropic>=1.0.0and a customAnthropicProviderhttp_clientmust be anhttpx2.AsyncClient(2.32.0/2.33.0). claiCLI (2.36.0):--mcp-configsupport and tool-call streaming. OpenRouter (2.30.0/2.32.0):openrouter:web_searchfor web search and web-search sources inprovider_details["annotations"].- Security (upgrade):
2.24.0fixes unbounded memory in the localweb_fetchtool /FileUrlmedia downloads (GHSA-v2xh-2vp8-57h8);2.27.1fixes a low-severity retry-prompt redaction leak wheninclude_content=False(GHSA-3gh4-cghq-f8v4);2.28.0fixes a high-severity missing content-type check on the dev web chat UI so arbitrary cross-origin requests no longer run the agent (GHSA-h4xc-3qfq-jf93);2.30.0fixes DNS-rebinding viaHost-header validation on the dev web chat UI, with an opt-inallowed_hostsfor non-loopback deployments (GHSA-q2xc-rrxj-58x9).
Release Highlights (2.13.0 -> 2.22.0)
- Durability capabilities replace wrapper agents:
TemporalDurability,DBOSDurability, andPrefectDurability(2.14.0) attach to a regularAgentviacapabilities=[...], replacing the deprecatedTemporalAgent/DBOSAgent/PrefectAgentwrapper classes (removed in v3). Existing wrapper-based workflows keep replaying correctly after switching. Seeintegrations.mdfor the updated Temporal/DBOS/Prefect examples. - New models/providers: Claude Opus 5 (
2.20.0),gemini-3.6-flash/gemini-3.5-flash-lite(2.16.0), Mistralreasoning_effort(2.14.0) andmistral_prompt_cache_key(2.16.0), OpenAI explicit prompt caching forgpt-5.6(2.15.0),BedrockMantleProvider(2.18.0), and theAdvisorToolbuiltin tool for Anthropic/OpenRouter (2.18.0). - Usage & limits:
cache_hit_ratioonRequestUsage/RunUsage(2.13.0),ToolFailedfor model-visible failures that don't consume a retry (2.16.0), optionalrun_id=on runs (2.16.0), tool-retry budget overrides atrun/iter/overridetime (2.15.0), andper_request_input_tokens_limitonUsageLimits(2.21.0). - Error handling & instrumentation:
ModelHTTPErrornow carriesheadersand a parsedretry_afterfrom every provider SDK (2.19.0);RaiseContentFilterErrorcapability andinclude_model_request_parametersinstrumentation setting (2.13.0); per-message OTel serialization is cached to avoidO(n^2)cost (2.17.0). - Dependencies: the
fastmcpoptional group now constrainsfastmcp<4(2.19.0).
Release Highlights (2.0.0 -> 2.12.0)
- V2 stable (
2.0.0): harness-first design with capabilities as the core primitive, a single composable unit bundling an agent's tools, hooks, instructions, and model settings. Migrating from V1 requires the official upgrade guide; the capability-based paths deprecated across the 1.9x line are now the default. - Message history & processing:
message_historyis provider-valid with repaired tool-call pairing,HistoryProcessoris exported, and deferred tool-call events plusEnqueuedMessagesEventare available (2.10.0to2.12.0). - New models/providers: Moonshot AI
kimi-k3, OpenAI background mode, and Anthropicstop_reason=pause_turnsupport; standardized reasoning-effort handling across Groq, DeepSeek, and Cerebras. - Embeddings & tokens: Gemini 2 text-prefix conditioning for embeddings and output audio-token accumulation in usage tracking.
- Fixes:
ToolReturnPartserialization uses field aliases, Anthropic/Bedrock native output schema handling, and actionable hints on usage-limit and tool-retry errors.
Release Highlights (1.96.1)
- V2 preparation:
Agent(..., prepare_tools=..., prepare_output_tools=..., event_stream_handler=...)is now the deprecated path; capability-based migration is the new direction. - Capability migration: use
PrepareTools,PrepareOutputTools, andProcessEventStreamcapabilities instead of wiring those behaviors through constructor sugar. - OpenAI fixes: the latest patch line also tightens
OpenAIResponsesModelsystem-prompt-role handling and image-generation request shaping.
Release Highlights (1.97.0 -> 1.102.0)
- MCP migration: prefer
MCPToolsetfor new MCP integrations.FastMCPToolsetand the olderMCPServer*client surface are now on the deprecation path. - Google provider split:
GoogleProviderandGoogleCloudProviderare now distinct, and model ids move fromgoogle-gla:/google-vertex:togoogle:/google-cloud:. - Streaming migration: move from
stream_responses()tostream_response(); the newer API yieldsModelResponseobjects directly. - Retry configuration: prefer
Agent(retries=...)orAgentRetries(...)over older constructor-level retry knobs. - New runtime tools:
ctx.enqueue()and MCP background tasks make it easier to queue follow-up work without forcing it into the current response turn.
Release Highlights (1.105.0 -> 1.107.0)
- New models: Claude Fable 5 and Claude Mythos 5 are supported (1.107.0), alongside Grok 4.3
reasoning_effortand updated xAI model names (1.105.0). - Deferred loading: instructions, tools, model settings, and hooks can now be loaded on demand instead of eagerly at agent construction (1.105.0).
- Model introspection:
known_model_names()enumeratesKnownModelNamemembers (1.107.0). - OpenRouter caching:
CachePointand prompt caching are implemented for OpenRouter (1.107.0). - xAI config:
XaiProvidergainsapi_hostandtimeout, plusseedparameter mapping (1.106.0). - Security:
VercelAIAdapterUploadedFilehandling was hardened against a confused-deputy file-read vulnerability (GHSA-h7p7-w5gc-xj3w, 1.106.0). - Fixes: incomplete streamed responses when
event_stream_handlerdoesn't consume the stream,from_data_urion non-base64 data URIs, Temporalgateway/model construction,GoogleModelSettings.google_cached_contentrequest shaping, Anthropic Bedrockmessage=Nonestart events, andAnthropicModel.count_tokenswith native tools.
Release Highlights (1.103.0 -> 1.104.0)
- MCP prompts: maintained
McpServerintegrations can now list and fetch prompts withlist_promptsandget_prompt; still preferMCPToolsetfor new client code unless legacy wrappers are required. - UI adapters:
VercelAIAdapterround-trips message timestamps throughUIMessage.metadata, andUIAdapter.sanitize_messagesstrips client-submittedforce_downloadfromFileUrlparts. - Provider/model updates: Claude Opus 4.8 is supported,
OpenRouterModelcan useanthropic_eager_input_streaming, and hybrid OpenRouter/xAI/Bedrock routes now forwardthinking=Falseconsistently. - Bedrock/toolset fixes: Bedrock maps malformed model/tool output to
FinishReason.error, recognizes adaptive thinking, preserves single-tooltool_choicecache behavior, and toolset prepare callbacks warn when they accidentally returnNone.
Release Highlights (1.75.0 -> 1.84.1)
- Capabilities:
CapabilityOrderingadds explicit wrapping/ordering control (innermost,outermost,wraps,wrapped_by,requires) for complex capability stacks. - Compaction: new server-side compaction capabilities for OpenAI and Anthropic; OpenAI adds stateful compaction mode.
- Models: Claude Opus 4.7 support and a native
OllamaModelpath with corrected Ollama capability flags for structured output. - Tools: tool hooks now consistently receive dict-shaped validated args for single-
BaseModeltools, and internal output tools skip hook execution. - Hardening: Google
FileSearchToolparsing received regex hardening in the1.83/1.84line.
Release Highlights (1.71.0 → 1.74.0)
- Capabilities: composable, reusable units of agent behavior that bundle tools, lifecycle hooks, instructions, and model settings into a single class. Plug into any agent for maximum reuse.
- AgentSpec: load agents from YAML/JSON files via
Agent.from_file. SupportsTemplateStrfor templated instructions referencing deps. - Hooks capability: define hooks using decorators (
@hooks.on_model_request, etc.). - Thinking capability: cross-provider
thinkingmodel setting for reasoning. - Provider-adaptive tools:
WebSearch,WebFetch,MCP,ImageGeneration— automatically fall back from builtin (provider) tools to local tools. - Online evaluation: evaluation infrastructure in
pydantic-evals. TextContent: user prompts withmetadatanot sent to model.- CaseLifecycle hooks: hooks for
Dataset.evaluatelifecycle. - Model swapping in hooks:
before_model_request/ wrap hooks can swap models viaModelRequestContext. ModelRetryfrom hooks: hooks can raiseModelRetryfor retry control flow.- Sync tool preparation functions supported.
MCPcapability no longer requires expliciturl=.
Release Highlights (1.69.0 → 1.70.0)
- Agents:
Agent(description=...)adds a human-readable description to the run span asgen_ai.agent.descriptionwhen instrumentation is enabled. - Models:
FallbackModelnow supports response-based fallback handlers for semantic failures in non-streaming runs. - Tools: multimodal tool results are passed directly to provider APIs instead of always being split into extra user-message parts.
- Bedrock:
bedrock_inference_profileis available on model and embedding settings for routing requests through an inference profile ARN. - Stability: provider fixes landed for OpenRouter Anthropic model matching, Cohere embeddings, Google image sizes, Bedrock tool-name sanitization, and malformed tool-call retry handling.
Quick Start
Basic Agent
from pydantic_ai import Agent
agent = Agent(
'openai:gpt-4o',
instructions='Be concise, reply with one sentence.'
)
result = agent.run_sync('Where does "hello world" come from?')
print(result.output)
With Structured Output
from pydantic import BaseModel
from pydantic_ai import Agent
class CityInfo(BaseModel):
name: str
country: str
population: int
agent = Agent('openai:gpt-4o', output_type=CityInfo)
result = agent.run_sync('Tell me about Paris')
print(result.output) # CityInfo(name='Paris', country='France', population=2161000)
With Tools and Dependencies
from dataclasses import dataclass
from pydantic_ai import Agent, RunContext
@dataclass
class Deps:
user_id: int
agent = Agent('openai:gpt-4o', deps_type=Deps)
@agent.tool
async def get_user_name(ctx: RunContext[Deps]) -> str:
"""Get the current user's name."""
return f"User #{ctx.deps.user_id}"
result = agent.run_sync('What is my name?', deps=Deps(user_id=123))
Key Features
| Feature | Description |
|---|---|
| Type-safe | Full IDE support, type checking |
| Model-agnostic | 30+ providers supported |
| Dependency Injection | Pass context to tools |
| Structured Output | Pydantic model validation |
| Embeddings | Multi-provider vector support |
| Logfire Integration | Built-in observability |
| MCP Support | External tools and data |
| Evals | Systematic testing |
| Graphs | Complex workflow support |
Supported Models
| Provider | Models |
|---|---|
| OpenAI | GPT-4o, GPT-4, o1, o3 |
| Anthropic | Claude Opus 5, Claude Opus 4.8, Claude 4, Claude 3.5 |
| Gemini 2.0, Gemini 1.5 | |
| xAI | Grok-4 (native SDK) |
| Groq | Llama, Mixtral |
| Mistral | Mistral Large, Codestral |
| Azure | Azure OpenAI |
| Bedrock | AWS Bedrock + Nova 2.0 |
| SambaNova | SambaNova models |
| Ollama | Local models |
Best Practices
- Use type hints — enables IDE support and validation
- Define output types — guarantees structured responses
- Use dependencies — inject context into tools
- Add tool docstrings — LLM uses them as descriptions
- Enable Logfire — for production observability
- Use
run_syncfor simple cases —runfor async - Override deps for testing —
agent.override(deps=...) - Set usage limits — prevent infinite loops with
UsageLimits
Prohibitions
- Do not expose API keys in code
- Do not skip output validation in production
- Do not ignore tool errors
- Do not use
run_streamwithout handling partial outputs - Do not forget to close MCP connections (
async with agent) - Do not assume capability order is arbitrary once multiple wrappers/hooks are involved; define it explicitly when composition matters.
Common Patterns
Streaming Response
async with agent.run_stream('Query') as response:
async for text in response.stream_text():
print(text, end='')
Fallback Models
from pydantic_ai.models.fallback import FallbackModel
fallback = FallbackModel(openai_model, anthropic_model)
agent = Agent(fallback)
MCP Integration
from pydantic_ai.mcp import MCPToolset
toolset = MCPToolset(command='python', args=['mcp_server.py'])
agent = Agent('openai:gpt-4o', toolsets=[toolset])
Testing with TestModel
from pydantic_ai.models.test import TestModel
agent = Agent(model=TestModel())
result = agent.run_sync('test') # Deterministic output
Embeddings
from pydantic_ai import Embedder
embedder = Embedder('openai:text-embedding-3-small')
# Embed search query
result = await embedder.embed_query('What is ML?')
# Embed documents for indexing
docs = ['Doc 1', 'Doc 2', 'Doc 3']
result = await embedder.embed_documents(docs)
See embeddings.md for providers and settings.
xAI Provider
from pydantic_ai import Agent
agent = Agent('xai:grok-4-1-fast-non-reasoning')
See models.md for configuration details.
Exa Neural Search
import os
from pydantic_ai import Agent
from pydantic_ai.common_tools.exa import ExaToolset
api_key = os.getenv('EXA_API_KEY')
toolset = ExaToolset(api_key, num_results=5, include_search=True)
agent = Agent('openai:gpt-4o', toolsets=[toolset])
See tools.md for all Exa tools.