Production Agent Checklist v1.0 发布
前 21 天,我一直在讲一个判断:
Agent Runtime 不是框架名。
它是一组工程责任。
今天把这组责任整理成第一版清单:
Production Agent Checklist v1.0
这不是最终版。
也不是'照着打勾就永远安全'的东西。
它的作用更朴素:让团队在讨论一个 Agent 能不能上线时,不再只说'我试了几次,感觉还行'。
感觉不够。
要有检查项。
这份清单适合谁
适合三类人:
正在做 coding agent / data agent / internal agent 的工程师
准备二开开源 agent framework 的团队
要评审 Agent 上线风险的技术负责人
不适合拿来做 PPT 装饰。
它应该进入:
README
PR 模板
release checklist
eval report
incident review
上线审批
如果一项检查永远不会影响发布决策,那它就不该留在清单里。
Production Agent Checklist v1.0
下面是第一版正文。
1. Run Identity
- 每次用户请求都有
runId、turnId、traceId。 - 每次模型决策或工具推进都有
stepId或 step index。 - 每个 tool call 有稳定
toolCallId。 - final、failed、cancelled、paused、max_steps、budget_exhausted 有明确 stop reason。
- 用户能看到任务是完成、失败、取消还是等待审批。
没有 ID,后面所有排障都是散的。
2. Loop Control
- Runtime 有 max steps。
- Runtime 有 wall-clock timeout。
- Runtime 能响应用户取消。
- Runtime 能识别 repeated tool call。
- Runtime 能识别 repeated error。
- Runtime 不允许模型无限重试同一个失败路径。
- 停止原因写入 trace 和最终响应。
Agent 要会做事,也要会停。

