Designed for
- Agent-platform engineers
- Evaluation and reliability teams
- MCP tool authors
- Teams reviewing failed automated workflows
Tetrees AI Pack · TAIP/1 · v1.0.0
A forensic evaluator for long agent runs that normalizes trajectories, derives invariants, locates the first critical failure, and separates root cause from downstream damage.
Tetrees Agent guide
Long agent runs fail through interacting plans, tools, policies, and observations; reviewing only the final answer confuses the first causal failure with later symptoms and produces vague remediation.
Start with one of these requests, then replace the details with your own.
A customer-service run ends with the wrong order change after several apparently reasonable tool calls.
Analyze this trajectory against the attached policy and tool schemas. Identify the first critical failure, distinguish downstream errors, classify it, and propose one regression case.
Expected outcome: A normalized trace, invariant violations with step evidence, earliest causal failure, taxonomy label, downstream chain, and a targeted regression test.
A new release reaches the right final answer but violates policy during execution.
Compare these two trajectories. Check plan-policy alignment, tool arguments, observation interpretation, recovery, and whether final-answer success hides an unsafe intermediate step.
Expected outcome: A step-aligned comparison, hidden policy violation, causal attribution, confidence, and release-blocking criteria.
Choose OpenAI, Claude or Z.AI in Agent Studio, then run with Tetrees Points or BYOK.
Hosted runs use only this signed set. Local MCP may add tools that you separately configure and approve.
Agent AVCP
テスト済み TAIP 成果物は本レポートと署名に結合されています。スコアは試験証拠であり将来の全出力を保証しません。
Overall score
9.5
out of 10
Mandatory gates
実行または取得にはログインしてください。
まだレビューがありません。