권장 사용자
- Agent-platform engineers
- Evaluation and reliability teams
- MCP tool authors
- Teams reviewing failed automated workflows
Tetrees AI Pack · TAIP/1 · v1.0.0
A forensic evaluator for long agent runs that normalizes trajectories, derives invariants, locates the first critical failure, and separates root cause from downstream damage.
Tetrees 에이전트 가이드
Long agent runs fail through interacting plans, tools, policies, and observations; reviewing only the final answer confuses the first causal failure with later symptoms and produces vague remediation.
아래 요청을 시작점으로 삼고 세부 내용을 상황에 맞게 바꾸세요.
A customer-service run ends with the wrong order change after several apparently reasonable tool calls.
Analyze this trajectory against the attached policy and tool schemas. Identify the first critical failure, distinguish downstream errors, classify it, and propose one regression case.
예상 결과: A normalized trace, invariant violations with step evidence, earliest causal failure, taxonomy label, downstream chain, and a targeted regression test.
A new release reaches the right final answer but violates policy during execution.
Compare these two trajectories. Check plan-policy alignment, tool arguments, observation interpretation, recovery, and whether final-answer success hides an unsafe intermediate step.
예상 결과: A step-aligned comparison, hidden policy violation, causal attribution, confidence, and release-blocking criteria.
Agent Studio에서 OpenAI, Claude 또는 Z.AI를 선택하고 Tetrees 포인트나 BYOK로 실행하세요.
호스팅 실행은 서명된 도구만 사용합니다. 로컬 MCP에서는 별도로 설정하고 승인한 도구를 추가할 수 있습니다.
Agent AVCP
테스트한 TAIP 아티팩트는 이 리포트와 플랫폼 서명에 결합됩니다. 점수는 테스트 근거이며 모든 향후 모델 응답을 보장하지 않습니다.
종합 점수
9.5
10점 만점
필수 게이트
이 팩을 실행하거나 추가하려면 로그인하세요.
아직 리뷰가 없습니다.