Agent Platform Eval Flywheel skill: what it does and how to install it
Evaluates and improves AI models and agents on Google Cloud with synthetic scenarios, rubric metrics, failure analysis and result comparison.
Summary generated from the skill's documentation.
Install
$ npx skills add google/skills --skill agent-platform-eval-flywheelRun it in a terminal. If your agent is already running, start a new session so it picks the skill up.
About this skill
What it does. Guides the Agent Platform GenAI Evaluation SDK through dataset preparation, inference, grading, failure analysis and iterative optimisation. It supports session traces, DataFrames, synthetic scenarios, built-in and custom metrics, persisted JSON and HTML reports, and before-and-after comparisons.
When to use it. For evaluating Google Cloud models or agents, selecting metrics, analysing failures or checking whether fixes improve results. It does not cover fine-tuning or general production deployment.
History
Repo stars
20.6kAbout +2.3k since 9 Jul 2026
Before 1 Oct 2026 the curve is estimated from public event data.
Stars are counted for the whole repository, which holds 36 skills.
Show as a table
| Date | Repo stars |
|---|---|
| 1 Oct 2026 | 20,557 |
| 24 Sept 2026 (estimated) | 20,498 |
| 17 Sept 2026 (estimated) | 20,452 |
| 10 Sept 2026 (estimated) | 19,878 |
| 3 Sept 2026 (estimated) | 18,833 |
| 27 Aug 2026 (estimated) | 18,776 |
| 20 Aug 2026 (estimated) | 18,765 |
| 13 Aug 2026 (estimated) | 18,765 |
| 6 Aug 2026 (estimated) | 18,294 |
| 30 Jul 2026 (estimated) | 18,271 |
| 23 Jul 2026 (estimated) | 18,271 |
| 16 Jul 2026 (estimated) | 18,260 |
| 9 Jul 2026 (estimated) | 18,248 |
Installs
7.1k
Tracking since . A chart appears once there are 7 days of data.
Installs via skills.sh
Similar skills
- Gemini APIGuides use of the Gemini API on Google Cloud's Agent Platform with the Google Gen AI SDK.
- gcloudAdds validation and guardrails to gcloud CLI operations across Google Cloud services, and trims their output.
- Skill CreatorCreates new skills, improves existing ones and measures how well they perform with evals and benchmarks.