适用对象
- Agent-platform engineers
- Evaluation and reliability teams
- MCP tool authors
- Teams reviewing failed automated workflows
Tetrees AI Pack · TAIP/1 · v1.0.0
A forensic evaluator for long agent runs that normalizes trajectories, derives invariants, locates the first critical failure, and separates root cause from downstream damage.
Tetrees Agent 指南
Long agent runs fail through interacting plans, tools, policies, and observations; reviewing only the final answer confuses the first causal failure with later symptoms and produces vague remediation.
可从以下请求开始,再替换为你的具体内容。
A customer-service run ends with the wrong order change after several apparently reasonable tool calls.
Analyze this trajectory against the attached policy and tool schemas. Identify the first critical failure, distinguish downstream errors, classify it, and propose one regression case.
预期结果: A normalized trace, invariant violations with step evidence, earliest causal failure, taxonomy label, downstream chain, and a targeted regression test.
A new release reaches the right final answer but violates policy during execution.
Compare these two trajectories. Check plan-policy alignment, tool arguments, observation interpretation, recovery, and whether final-answer success hides an unsafe intermediate step.
预期结果: A step-aligned comparison, hidden policy violation, causal attribution, confidence, and release-blocking criteria.
在 Agent Studio 中选择 OpenAI、Claude 或 Z.AI,然后使用 Tetrees 积分或 BYOK 运行。
托管运行仅使用此签名工具集。本地 MCP 可添加由你单独配置和批准的工具。
Agent AVCP
测试的 TAIP 制品与本报告及平台签名绑定。分数反映测试证据,不保证未来每次模型输出。
总分
9.5
满分 10 分
必需门槛
请登录后运行或获取此包。
暂无评价。