Pydantic AI 拆解:Agent 容器、Capability 组合与一个银行客服 agent
Pydantic AI 是 Pydantic 团队做的 Python AI SDK / agent 框架,定位是一个类型安全的 agent loop,模型可以一行字符串替换。它想把 FastAPI 的开发手感搬到 GenAI/Agent 开发里,要求 Python 3.10+。Pydantic Validation 本身就是 OpenAI SDK、Anthropic SDK、Google ADK、LangChain 的验证层,也是 FastAPI 的基础,这套类型约束现在被直接搬到了 agent 上。
- 项目地址:https://github.com/pydantic/pydantic-ai
- 文档:https://pydantic.dev/docs/ai/
- 官网:https://pydantic.dev/pydantic-ai
Agent 是一个容器
Agent 是与 LLM 交互的主要接口,可以理解成一个容器,装着 instructions(系统提示)、tools、hooks、model settings 和 model。它的复用方式和 FastAPI App 一样:可以全局实例化一个到处用,也可以动态创建任意多个。默认泛型是 Agent[object, str],即依赖类型为 object、输出类型为 str。
运行方式有 5 种:
agent.run()—— async,返回RunResultagent.run_sync()—— 同步版本,内部就是loop.run_until_complete(self.run())agent.run_stream()—— async context manager,返回StreamedRunResult,可流式拿到文本和结构化输出agent.run_stream_events()—— 拿到AgentStreamEvent的 async 迭代器,最后一个AgentRunResultEvent带最终结果agent.iter()—— 返回AgentRun,可对 agent 底层 Graph 的节点做 async 迭代
底层执行流由 pydantic-graph 管理,那是一个泛型的、以类型为中心的有限状态机库,可以脱离 Pydantic AI 单独使用。
结构化输出与校验重试
给 agent 一个 output_type(Pydantic 模型),每次 run 都返回校验过、带类型的对象。Pydantic 会据此构建 JSON Schema 告诉 LLM 怎么返回数据,run 结束时做校验;校验失败,agent 会被提示重试。
from typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
class Sentiment(BaseModel):
label: Literal['positive', 'negative', 'neutral']
score: float = Field(ge=-1, le=1)
agent = Agent('openai:gpt-6-sol', output_type=Sentiment)
@agent.tool
def recent_reviews(ctx: RunContext, product: str) -> list[str]:
"""Fetch recent review snippets for a product."""
return ['The new release fixed everything I complained about!']
result = agent.run_sync('How are people feeling about the Extract app?')
print(result.output)
#> label='positive' score=0.9
@agent.tool 函数接收 RunContext 携带依赖;其余签名和 docstring 会变成给 LLM 的 tool schema,参数在你的代码运行前就已被校验。
依赖注入
Pydantic AI 提供可选的依赖注入系统,用来给 agent 的 instructions、tools、output validators 提供数据和服务。依赖通过 RunContext 参数传递,类型由 deps_type 决定;类型注解写错,静态类型检查器会直接抓到。这套机制在做单元测试和 eval 驱动的迭代开发时特别有用。
Capability:可组合的单一原语
capability(能力)是 Pydantic AI 的一个核心抽象:把 tools、instructions、hooks、model settings 打包成可复用的单元。core 自带 MCP、web search 等基础能力,Harness(官方能力库)提供其余部分,像 Coder、Researcher 这样的完整 agent 本身就是 capability 组合出来的,拆开和拼起来一样简单。也可以用 YAML/JSON 定义 agent spec,完全不写代码。
一个完整 coding agent 示例,在终端里跑:
uv add pydantic-ai pydantic-ai-harness
from pydantic_ai import Agent
from pydantic_ai.capabilities import WebSearch
from pydantic_ai_harness import Advisor, Coder
agent = Agent(
'anthropic:claude-fable-5-1',
capabilities=[
Coder(), # files, shell, repo context, sub-agents, context management
WebSearch(),
Advisor('openai:gpt-6-sol'), # 卡住时让另一个模型给第二意见
],
)
agent.to_cli_sync()
Coder 本身也是一个普通的组合 capability,等价于这些块:
capabilities = [
FileSystem('.'), Shell(cwd='.'), RepoContext(), SubAgents(...),
ClearToolResults(), WarnNearLimits(), ToolOutputLimits(), RepairToolArguments(),
]
银行客服 agent 完整示例
这个例子把依赖注入、function tools、结构化输出、可复用 capability 以及按需加载(on-demand)capability 都串起来了:
from dataclasses import dataclass
from pydantic import BaseModel, Field
from pydantic_ai import Agent, Capability, RunContext
from bank_database import DatabaseConn
@dataclass
class SupportDependencies: # 注入任意 client:DB 连接池、HTTP API、用户信息
customer_id: int
db: DatabaseConn
class SupportOutput(BaseModel):
support_advice: str = Field(description='Advice returned to the customer')
block_card: bool = Field(description="Whether to block the customer's card")
risk: int = Field(description='Risk level of query', ge=0, le=10)
customer_context = Capability[SupportDependencies](
id='customer-context',
description="Who the customer is and what's on their account.",
)
@customer_context.instructions
async def add_customer_name(ctx: RunContext[SupportDependencies]) -> str:
customer_name = await ctx.deps.db.customer_name(id=ctx.deps.customer_id)
return f"The customer's name is {customer_name!r}"
@customer_context.tool
async def customer_balance(
ctx: RunContext[SupportDependencies], include_pending: bool
) -> float:
"""Returns the customer's current account balance."""
return await ctx.deps.db.customer_balance(
id=ctx.deps.customer_id, include_pending=include_pending,
)
refunds = Capability[SupportDependencies](
id='refunds', description='Refund eligibility and refund status.', defer_loading=True,
)
@refunds.tool
async def refund_status(ctx: RunContext[SupportDependencies]) -> str:
"""Look up the refund status for the customer's most recent charge."""
return await ctx.deps.db.refund_status(id=ctx.deps.customer_id)
support_agent = Agent(
'openai:gpt-6-sol',
deps_type=SupportDependencies,
output_type=SupportOutput,
instructions=(
'You are a support agent in our bank, give the '
'customer support and judge the risk level of their query.'
),
capabilities=[customer_context, refunds],
)
注意 refunds 上的 defer_loading=True:这组工具不是一开始就塞进上下文,而是按需加载。
可观测性
日志和追踪走标准 OpenTelemetry,任何 OTel 后端都能接。最省事的是 Logfire,两行:
import logfire
logfire.configure()
logfire.instrument_pydantic_ai()
instrument 一行开启后,每个 model call 和 tool call 都会出现在 Logfire / OTel 后端。
持久执行
给 agent 挂上 TemporalDurability,同一个 agent 就能跑在 Temporal workflow 里,用持久执行。每次 model/tool 调用变成 durable activity,跑后台队列的 run 能撑过重启、失败和长时间等待。DBOS、Prefect 也是同样方式接入,另外还有 Restate、AWS Lambda、Kitaru、Airflow、Absurd,一共八个引擎。
uv add "pydantic-ai[temporal]"
模型支持与测试
模型和 provider 覆盖得比较全:OpenAI、Anthropic、Google、Bedrock、Azure AI Foundry、Groq、Mistral、xAI、Ollama、DeepSeek、Cohere 等,一行字符串即可切换,也可以通过 Pydantic AI Gateway(一个 key 覆盖全部,带 failover 和成本监控)。
测试侧,Pydantic Evals 像 pytest 测代码一样测 agent 行为;内置的 test model(TestModel)不需要 API key 就能起步。
从别的框架迁移的话,comparisons 文档列了 Pydantic AI 与 LangChain、Google ADK、Claude Agent SDK 等十来个框架的差异,migration skills 可以让 coding agent 帮你把已有应用迁过来。
持久执行那一块我只实际跑通了 Temporal,DBOS、Prefect 和其余几个引擎的接入成本还没横向比过;defer_loading 的收益也取决于工具规模,工具少时开不开差别不大。