対象ユーザー
- Agent-platform engineers
- Evaluation and reliability teams
- MCP tool authors
- Teams reviewing failed automated workflows
Tetrees AI Pack · TAIP/1 · v1.0.0
A forensic evaluator for long agent runs that normalizes trajectories, derives invariants, locates the first critical failure, and separates root cause from downstream damage.
Tetrees Agent ガイド
Long agent runs fail through interacting plans, tools, policies, and observations; reviewing only the final answer confuses the first causal failure with later symptoms and produces vague remediation.
以下のリクエストを出発点にして、詳細を用途に合わせて置き換えてください。
A customer-service run ends with the wrong order change after several apparently reasonable tool calls.
Analyze this trajectory against the attached policy and tool schemas. Identify the first critical failure, distinguish downstream errors, classify it, and propose one regression case.
期待される結果: A normalized trace, invariant violations with step evidence, earliest causal failure, taxonomy label, downstream chain, and a targeted regression test.
A new release reaches the right final answer but violates policy during execution.
Compare these two trajectories. Check plan-policy alignment, tool arguments, observation interpretation, recovery, and whether final-answer success hides an unsafe intermediate step.
期待される結果: A step-aligned comparison, hidden policy violation, causal attribution, confidence, and release-blocking criteria.
Agent Studio で OpenAI、Claude、Z.AI を選び、Tetrees ポイントまたは BYOK で実行します。
ホスト実行では署名済みのツールだけを使用します。ローカル MCP では、別途設定して承認したツールを追加できます。
Agent AVCP
テスト済み TAIP 成果物は本レポートと署名に結合されています。スコアは試験証拠であり将来の全出力を保証しません。
総合スコア
9.5
10点満点
必須ゲート
実行または取得にはログインしてください。
まだレビューがありません。