GKE Inference skill: what it does and how to install it
Deploys AI/ML inference workloads on Google Kubernetes Engine using model servers, GPUs or TPUs, and generated Kubernetes manifests.
Summary generated from the skill's documentation.
Install
$ npx skills add google/skills --skill gke-inferenceRun it in a terminal. If your agent is already running, start a new session so it picks the skill up.
About this skill
What it does. Uses Google Cloud's inference profiles to find supported model and accelerator combinations, generate Kubernetes manifests, and guide deployment on GKE. Covers accelerator selection, GPU ComputeClasses, autoscaling, performance tuning and troubleshooting.
When to use it. For deploying and optimising AI model serving on GKE, including LLM inference. It is not for migrating existing AI workloads, GKE RAG with Cloud SQL or AlloyDB, or batch and HPC workloads.
History
Repo stars
20.9kAbout +2.7k since 9 Jul 2026
Before 1 Oct 2026 the curve is estimated from public event data.
Stars are counted for the whole repository, which holds 61 skills.
Show as a table
| Date | Repo stars |
|---|---|
| 5 Oct 2026 | 20,939 |
| 4 Oct 2026 | 20,924 |
| 3 Oct 2026 | 20,865 |
| 2 Oct 2026 | 20,592 |
| 1 Oct 2026 | 20,557 |
| 24 Sept 2026 (estimated) | 20,498 |
| 17 Sept 2026 (estimated) | 20,452 |
| 10 Sept 2026 (estimated) | 19,878 |
| 3 Sept 2026 (estimated) | 18,833 |
| 27 Aug 2026 (estimated) | 18,776 |
| 20 Aug 2026 (estimated) | 18,765 |
| 13 Aug 2026 (estimated) | 18,765 |
| 6 Aug 2026 (estimated) | 18,294 |
| 30 Jul 2026 (estimated) | 18,271 |
| 23 Jul 2026 (estimated) | 18,271 |
| 16 Jul 2026 (estimated) | 18,260 |
| 9 Jul 2026 (estimated) | 18,248 |
Installs
5.6k
Tracking since . A chart appears once there are 7 days of data.
Installs via skills.sh
Similar skills
- Using Git WorktreesSets up an isolated workspace with git worktrees before feature work or plan execution begins.
- Domain ModelingBuilds and sharpens a project's domain model through a glossary and ADRs.
- PrototypeBuilds a throwaway prototype to answer a design question about a state model, logic or UI.