
#57 by installsTesting & QA
google-agents-cli-eval
Run and iterate agent evaluations with datasets, metrics, and LLM-as-judge scoring
473K
53.0K
6.0K
Apr 21, 2026
Install
npx skills add https://github.com/google/agents-cli --skill google-agents-cli-evalWhat it does
Guides you through the Agent Platform eval methodology—creating datasets, configuring metrics, analyzing failures, and optimizing agent quality. Reach for this when you need to measure agent performance, understand what's breaking, or compare results across iterations. Works with any agents-cli project regardless of framework.
- Execute eval runs with agents-cli: generate synthetic data, grade against metrics, get scored results
- Design evaluation datasets and custom metrics; reference canonical schema and built-in judge models
- Analyze eval failures, compare results across runs, and optimize agent behavior iteratively
- Handle multimodal, live/voice, and multi-turn agent evals with framework-specific patterns