Dành cho
- Agent-platform engineers
- Evaluation and reliability teams
- MCP tool authors
- Teams reviewing failed automated workflows
Tetrees AI Pack · TAIP/1 · v1.0.0
A forensic evaluator for long agent runs that normalizes trajectories, derives invariants, locates the first critical failure, and separates root cause from downstream damage.
Hướng dẫn Tetrees Agent
Long agent runs fail through interacting plans, tools, policies, and observations; reviewing only the final answer confuses the first causal failure with later symptoms and produces vague remediation.
Bắt đầu với một yêu cầu bên dưới rồi thay chi tiết bằng dữ liệu của bạn.
A customer-service run ends with the wrong order change after several apparently reasonable tool calls.
Analyze this trajectory against the attached policy and tool schemas. Identify the first critical failure, distinguish downstream errors, classify it, and propose one regression case.
Kết quả dự kiến: A normalized trace, invariant violations with step evidence, earliest causal failure, taxonomy label, downstream chain, and a targeted regression test.
A new release reaches the right final answer but violates policy during execution.
Compare these two trajectories. Check plan-policy alignment, tool arguments, observation interpretation, recovery, and whether final-answer success hides an unsafe intermediate step.
Kết quả dự kiến: A step-aligned comparison, hidden policy violation, causal attribution, confidence, and release-blocking criteria.
Chọn OpenAI, Claude hoặc Z.AI trong Agent Studio rồi chạy bằng Điểm Tetrees hoặc BYOK.
Chạy trên máy chủ chỉ dùng bộ công cụ đã ký này. MCP cục bộ có thể thêm công cụ do bạn tự cấu hình và phê duyệt.
Agent AVCP
TAIP đã kiểm tra được liên kết với báo cáo và chữ ký nền tảng. Điểm số là bằng chứng kiểm thử, không bảo đảm mọi phản hồi tương lai.
Điểm tổng
9.5
trên 10
Cổng bắt buộc
Đăng nhập để chạy hoặc nhận gói.
Chưa có đánh giá.