pydantic-ai

v2026.09.24

Pydantic AI Python agent framework. Covers typed tools, model providers, evals, MCP, UI adapters, and observability. Use when building Python AI agents with Pydantic AI, configuring model providers, implementing typed tools/dependencies, running evals, or integrating MCP servers. Keywords: pydantic-ai, agents, evals, MCP, Logfire.

GitHub
Install command
npx skhub add itechmeat/pydantic-ai
Markdown
SKILL.md

Pydantic AI

Python agent framework for building production-grade GenAI applications with the "FastAPI feeling".

Quick Navigation

TopicReference
Agentsagents.md
Capabilitiesagents.md
Toolstools.md
Modelsmodels.md
Embeddingsembeddings.md
Evalsevals.md
Integrationsintegrations.md
Graphsgraphs.md
UI Streamsui.md
Installationinstallation.md

When to Use

  • Building AI agents with structured output
  • Need type-safe, IDE-friendly agent development
  • Require dependency injection for tools
  • Multi-model support (OpenAI, Anthropic, Gemini, etc.)
  • Production observability with Logfire
  • Complex workflows with graphs

Installation

See references/installation.md for full/slim install options and optional dependency groups. Requires Python 3.10+.

Release Highlights (2.40.0 -> 2.48.0)

  • Security (upgrade) (2.44.0, backported to 1.107.6): four fixes reached through web_fetch_tool or OTel instrumentation — IPv6 zone-identifier bypass of the private-IP blocklist when local URLs are opted in, superlinear web_fetch HTML/charset processing that could stall the event loop, domain blocklists compared without resolver normalization, and span content leaks with include_content=False (GHSA-vmxc-h2x2-jmf3, GHSA-fpf4-vwcp-v4hp, GHSA-22h6-qm39-v87j, GHSA-4x9p-g9wm-8q7f).
  • Realtime session control (2.40.0/2.46.0): handle_barge_in=True with interrupt(played_bytes=...), RealtimeSession.enqueue() for out-of-band prompts, respond= on send(), wait_for_playback(), and @agent.on_event listeners.
  • Image generation (2.41.0): direct ImageGenerator API alongside the built-in tool; fallback_model on ImageGeneration/XSearch deprecated in favor of fallback_subagent_model.
  • Providers and models (2.41.0/2.42.0/2.48.0): openai-codex provider for ChatGPT/Codex subscription auth, GitHubCopilotProvider; OpenAI gpt-6-sol/gpt-6-luna and Claude Opus 5.5 (claude-opus-5-5).
  • Pricing (2.40.0): pydantic_ai.prices.update_in_background() refreshes genai-prices data in the background.
  • Temporal (2.46.0): event_stream_topic on TemporalDurability streams agent events via Workflow Streams.

Release Highlights (2.23.0 -> 2.39.0)

  • Realtime speech-to-speech (2.28.0): Agent.realtime() with browser WebRTC plus server sideband support; Azure AI Voice Live via the azure_voice_live setting (2.29.0); RealtimeSession.send_audio() accepts async iterables (2.36.0).
  • Run cancellation (2.26.0): AgentRun.cancel(), RunContext.cancel(), and RunCancelled; stream_run_events() returns a public AgentRunEvents handle with cancel() and run-state access.
  • Deferred tool reveal (2.23.0/2.26.0/2.30.0): ToolAvailabilityDeltaPart with native tool_addition/additional_tools; tools can stay hidden until revealed via tool search, load_capability, or ToolReturn.tools, on each provider's native deferral/addition channel; a deferred tool must be revealed before it can be called.
  • Cost tracking (2.23.0): cost on RunUsage and cost_limit on UsageLimits. Context window (2.38.0): context_window on ModelProfile and context_window_used on RunContext.
  • Typed events (2.38.0): application code and capabilities can emit typed CustomEvent/CapabilityEvent into the run event stream and subscribe with @on_event. Durable execution (2.36.0): @durable_operation for capabilities plus a public backend API for third-party durable execution engines.
  • Providers and models: Crusoe (2.28.0), Snowflake Cortex (2.27.0), and vLLM (2.38.0) providers; gemini-3.7-flash (2.30.0), gemini-3.8-flash (2.38.0), Claude Fable 5.1 / Mythos 5.1 (2.38.0), OpenAI gpt-6-astra (2.39.0), GLM-5.3 on ZaiModel (2.34.0), DeepSeek V4 Flash (2.26.0).
  • MCP and HTTP clients: MCPToolset supports FastMCP 4 and MCP SDK v2 alongside FastMCP 3 (2.29.0). Builds moved to httpx2 clients; pydantic-ai[anthropic] requires anthropic>=1.0.0 and a custom AnthropicProvider http_client must be an httpx2.AsyncClient (2.32.0/2.33.0).
  • clai CLI (2.36.0): --mcp-config support and tool-call streaming. OpenRouter (2.30.0/2.32.0): openrouter:web_search for web search and web-search sources in provider_details["annotations"].
  • Security (upgrade): 2.24.0 fixes unbounded memory in the local web_fetch tool / FileUrl media downloads (GHSA-v2xh-2vp8-57h8); 2.27.1 fixes a low-severity retry-prompt redaction leak when include_content=False (GHSA-3gh4-cghq-f8v4); 2.28.0 fixes a high-severity missing content-type check on the dev web chat UI so arbitrary cross-origin requests no longer run the agent (GHSA-h4xc-3qfq-jf93); 2.30.0 fixes DNS-rebinding via Host-header validation on the dev web chat UI, with an opt-in allowed_hosts for non-loopback deployments (GHSA-q2xc-rrxj-58x9).

Release Highlights (2.13.0 -> 2.22.0)

  • Durability capabilities replace wrapper agents: TemporalDurability, DBOSDurability, and PrefectDurability (2.14.0) attach to a regular Agent via capabilities=[...], replacing the deprecated TemporalAgent / DBOSAgent / PrefectAgent wrapper classes (removed in v3). Existing wrapper-based workflows keep replaying correctly after switching. See integrations.md for the updated Temporal/DBOS/Prefect examples.
  • New models/providers: Claude Opus 5 (2.20.0), gemini-3.6-flash/gemini-3.5-flash-lite (2.16.0), Mistral reasoning_effort (2.14.0) and mistral_prompt_cache_key (2.16.0), OpenAI explicit prompt caching for gpt-5.6 (2.15.0), BedrockMantleProvider (2.18.0), and the AdvisorTool builtin tool for Anthropic/OpenRouter (2.18.0).
  • Usage & limits: cache_hit_ratio on RequestUsage/RunUsage (2.13.0), ToolFailed for model-visible failures that don't consume a retry (2.16.0), optional run_id= on runs (2.16.0), tool-retry budget overrides at run/iter/override time (2.15.0), and per_request_input_tokens_limit on UsageLimits (2.21.0).
  • Error handling & instrumentation: ModelHTTPError now carries headers and a parsed retry_after from every provider SDK (2.19.0); RaiseContentFilterError capability and include_model_request_parameters instrumentation setting (2.13.0); per-message OTel serialization is cached to avoid O(n^2) cost (2.17.0).
  • Dependencies: the fastmcp optional group now constrains fastmcp<4 (2.19.0).

Release Highlights (2.0.0 -> 2.12.0)

  • V2 stable (2.0.0): harness-first design with capabilities as the core primitive, a single composable unit bundling an agent's tools, hooks, instructions, and model settings. Migrating from V1 requires the official upgrade guide; the capability-based paths deprecated across the 1.9x line are now the default.
  • Message history & processing: message_history is provider-valid with repaired tool-call pairing, HistoryProcessor is exported, and deferred tool-call events plus EnqueuedMessagesEvent are available (2.10.0 to 2.12.0).
  • New models/providers: Moonshot AI kimi-k3, OpenAI background mode, and Anthropic stop_reason=pause_turn support; standardized reasoning-effort handling across Groq, DeepSeek, and Cerebras.
  • Embeddings & tokens: Gemini 2 text-prefix conditioning for embeddings and output audio-token accumulation in usage tracking.
  • Fixes: ToolReturnPart serialization uses field aliases, Anthropic/Bedrock native output schema handling, and actionable hints on usage-limit and tool-retry errors.

Release Highlights (1.96.1)

  • V2 preparation: Agent(..., prepare_tools=..., prepare_output_tools=..., event_stream_handler=...) is now the deprecated path; capability-based migration is the new direction.
  • Capability migration: use PrepareTools, PrepareOutputTools, and ProcessEventStream capabilities instead of wiring those behaviors through constructor sugar.
  • OpenAI fixes: the latest patch line also tightens OpenAIResponsesModel system-prompt-role handling and image-generation request shaping.

Release Highlights (1.97.0 -> 1.102.0)

  • MCP migration: prefer MCPToolset for new MCP integrations. FastMCPToolset and the older MCPServer* client surface are now on the deprecation path.
  • Google provider split: GoogleProvider and GoogleCloudProvider are now distinct, and model ids move from google-gla: / google-vertex: to google: / google-cloud:.
  • Streaming migration: move from stream_responses() to stream_response(); the newer API yields ModelResponse objects directly.
  • Retry configuration: prefer Agent(retries=...) or AgentRetries(...) over older constructor-level retry knobs.
  • New runtime tools: ctx.enqueue() and MCP background tasks make it easier to queue follow-up work without forcing it into the current response turn.

Release Highlights (1.105.0 -> 1.107.0)

  • New models: Claude Fable 5 and Claude Mythos 5 are supported (1.107.0), alongside Grok 4.3 reasoning_effort and updated xAI model names (1.105.0).
  • Deferred loading: instructions, tools, model settings, and hooks can now be loaded on demand instead of eagerly at agent construction (1.105.0).
  • Model introspection: known_model_names() enumerates KnownModelName members (1.107.0).
  • OpenRouter caching: CachePoint and prompt caching are implemented for OpenRouter (1.107.0).
  • xAI config: XaiProvider gains api_host and timeout, plus seed parameter mapping (1.106.0).
  • Security: VercelAIAdapter UploadedFile handling was hardened against a confused-deputy file-read vulnerability (GHSA-h7p7-w5gc-xj3w, 1.106.0).
  • Fixes: incomplete streamed responses when event_stream_handler doesn't consume the stream, from_data_uri on non-base64 data URIs, Temporal gateway/ model construction, GoogleModelSettings.google_cached_content request shaping, Anthropic Bedrock message=None start events, and AnthropicModel.count_tokens with native tools.

Release Highlights (1.103.0 -> 1.104.0)

  • MCP prompts: maintained McpServer integrations can now list and fetch prompts with list_prompts and get_prompt; still prefer MCPToolset for new client code unless legacy wrappers are required.
  • UI adapters: VercelAIAdapter round-trips message timestamps through UIMessage.metadata, and UIAdapter.sanitize_messages strips client-submitted force_download from FileUrl parts.
  • Provider/model updates: Claude Opus 4.8 is supported, OpenRouterModel can use anthropic_eager_input_streaming, and hybrid OpenRouter/xAI/Bedrock routes now forward thinking=False consistently.
  • Bedrock/toolset fixes: Bedrock maps malformed model/tool output to FinishReason.error, recognizes adaptive thinking, preserves single-tool tool_choice cache behavior, and toolset prepare callbacks warn when they accidentally return None.

Release Highlights (1.75.0 -> 1.84.1)

  • Capabilities: CapabilityOrdering adds explicit wrapping/ordering control (innermost, outermost, wraps, wrapped_by, requires) for complex capability stacks.
  • Compaction: new server-side compaction capabilities for OpenAI and Anthropic; OpenAI adds stateful compaction mode.
  • Models: Claude Opus 4.7 support and a native OllamaModel path with corrected Ollama capability flags for structured output.
  • Tools: tool hooks now consistently receive dict-shaped validated args for single-BaseModel tools, and internal output tools skip hook execution.
  • Hardening: Google FileSearchTool parsing received regex hardening in the 1.83/1.84 line.

Release Highlights (1.71.0 → 1.74.0)

  • Capabilities: composable, reusable units of agent behavior that bundle tools, lifecycle hooks, instructions, and model settings into a single class. Plug into any agent for maximum reuse.
  • AgentSpec: load agents from YAML/JSON files via Agent.from_file. Supports TemplateStr for templated instructions referencing deps.
  • Hooks capability: define hooks using decorators (@hooks.on_model_request, etc.).
  • Thinking capability: cross-provider thinking model setting for reasoning.
  • Provider-adaptive tools: WebSearch, WebFetch, MCP, ImageGeneration — automatically fall back from builtin (provider) tools to local tools.
  • Online evaluation: evaluation infrastructure in pydantic-evals.
  • TextContent: user prompts with metadata not sent to model.
  • CaseLifecycle hooks: hooks for Dataset.evaluate lifecycle.
  • Model swapping in hooks: before_model_request / wrap hooks can swap models via ModelRequestContext.
  • ModelRetry from hooks: hooks can raise ModelRetry for retry control flow.
  • Sync tool preparation functions supported. MCP capability no longer requires explicit url=.

Release Highlights (1.69.0 → 1.70.0)

  • Agents: Agent(description=...) adds a human-readable description to the run span as gen_ai.agent.description when instrumentation is enabled.
  • Models: FallbackModel now supports response-based fallback handlers for semantic failures in non-streaming runs.
  • Tools: multimodal tool results are passed directly to provider APIs instead of always being split into extra user-message parts.
  • Bedrock: bedrock_inference_profile is available on model and embedding settings for routing requests through an inference profile ARN.
  • Stability: provider fixes landed for OpenRouter Anthropic model matching, Cohere embeddings, Google image sizes, Bedrock tool-name sanitization, and malformed tool-call retry handling.

Quick Start

Basic Agent

from pydantic_ai import Agent

agent = Agent(
    'openai:gpt-4o',
    instructions='Be concise, reply with one sentence.'
)

result = agent.run_sync('Where does "hello world" come from?')
print(result.output)

With Structured Output

from pydantic import BaseModel
from pydantic_ai import Agent

class CityInfo(BaseModel):
    name: str
    country: str
    population: int

agent = Agent('openai:gpt-4o', output_type=CityInfo)
result = agent.run_sync('Tell me about Paris')
print(result.output)  # CityInfo(name='Paris', country='France', population=2161000)

With Tools and Dependencies

from dataclasses import dataclass
from pydantic_ai import Agent, RunContext

@dataclass
class Deps:
    user_id: int

agent = Agent('openai:gpt-4o', deps_type=Deps)

@agent.tool
async def get_user_name(ctx: RunContext[Deps]) -> str:
    """Get the current user's name."""
    return f"User #{ctx.deps.user_id}"

result = agent.run_sync('What is my name?', deps=Deps(user_id=123))

Key Features

FeatureDescription
Type-safeFull IDE support, type checking
Model-agnostic30+ providers supported
Dependency InjectionPass context to tools
Structured OutputPydantic model validation
EmbeddingsMulti-provider vector support
Logfire IntegrationBuilt-in observability
MCP SupportExternal tools and data
EvalsSystematic testing
GraphsComplex workflow support

Supported Models

ProviderModels
OpenAIGPT-4o, GPT-4, o1, o3
AnthropicClaude Opus 5, Claude Opus 4.8, Claude 4, Claude 3.5
GoogleGemini 2.0, Gemini 1.5
xAIGrok-4 (native SDK)
GroqLlama, Mixtral
MistralMistral Large, Codestral
AzureAzure OpenAI
BedrockAWS Bedrock + Nova 2.0
SambaNovaSambaNova models
OllamaLocal models

Best Practices

  1. Use type hints — enables IDE support and validation
  2. Define output types — guarantees structured responses
  3. Use dependencies — inject context into tools
  4. Add tool docstrings — LLM uses them as descriptions
  5. Enable Logfire — for production observability
  6. Use run_sync for simple cases — run for async
  7. Override deps for testing — agent.override(deps=...)
  8. Set usage limits — prevent infinite loops with UsageLimits

Prohibitions

  • Do not expose API keys in code
  • Do not skip output validation in production
  • Do not ignore tool errors
  • Do not use run_stream without handling partial outputs
  • Do not forget to close MCP connections (async with agent)
  • Do not assume capability order is arbitrary once multiple wrappers/hooks are involved; define it explicitly when composition matters.

Common Patterns

Streaming Response

async with agent.run_stream('Query') as response:
    async for text in response.stream_text():
        print(text, end='')

Fallback Models

from pydantic_ai.models.fallback import FallbackModel

fallback = FallbackModel(openai_model, anthropic_model)
agent = Agent(fallback)

MCP Integration

from pydantic_ai.mcp import MCPToolset

toolset = MCPToolset(command='python', args=['mcp_server.py'])
agent = Agent('openai:gpt-4o', toolsets=[toolset])

Testing with TestModel

from pydantic_ai.models.test import TestModel

agent = Agent(model=TestModel())
result = agent.run_sync('test')  # Deterministic output

Embeddings

from pydantic_ai import Embedder

embedder = Embedder('openai:text-embedding-3-small')

# Embed search query
result = await embedder.embed_query('What is ML?')

# Embed documents for indexing
docs = ['Doc 1', 'Doc 2', 'Doc 3']
result = await embedder.embed_documents(docs)

See embeddings.md for providers and settings.

xAI Provider

from pydantic_ai import Agent

agent = Agent('xai:grok-4-1-fast-non-reasoning')

See models.md for configuration details.

Exa Neural Search

import os
from pydantic_ai import Agent
from pydantic_ai.common_tools.exa import ExaToolset

api_key = os.getenv('EXA_API_KEY')
toolset = ExaToolset(api_key, num_results=5, include_search=True)
agent = Agent('openai:gpt-4o', toolsets=[toolset])

See tools.md for all Exa tools.

Links

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/pydantic-ai

Default branch

master

Latest commit

7ae8a00

Tree SHA

47f5439