Skip to main content
gke-inferenceby GoogleOfficialDevelopmentGitHub stars: 20.9k

GKE Inference skill: what it does and how to install it

Deploys AI/ML inference workloads on Google Kubernetes Engine using model servers, GPUs or TPUs, and generated Kubernetes manifests.

Summary generated from the skill's documentation.

Install

$ npx skills add google/skills --skill gke-inference

Run it in a terminal. If your agent is already running, start a new session so it picks the skill up.

About this skill

What it does. Uses Google Cloud's inference profiles to find supported model and accelerator combinations, generate Kubernetes manifests, and guide deployment on GKE. Covers accelerator selection, GPU ComputeClasses, autoscaling, performance tuning and troubleshooting.

When to use it. For deploying and optimising AI model serving on GKE, including LLM inference. It is not for migrating existing AI workloads, GKE RAG with Cloud SQL or AlloyDB, or batch and HPC workloads.

History

Repo stars

20.9kAbout +2.7k since 9 Jul 2026

Before 1 Oct 2026 the curve is estimated from public event data.

Stars are counted for the whole repository, which holds 61 skills.

Show as a table
Repo stars by day
DateRepo stars
5 Oct 202620,939
4 Oct 202620,924
3 Oct 202620,865
2 Oct 202620,592
1 Oct 202620,557
24 Sept 2026 (estimated)20,498
17 Sept 2026 (estimated)20,452
10 Sept 2026 (estimated)19,878
3 Sept 2026 (estimated)18,833
27 Aug 2026 (estimated)18,776
20 Aug 2026 (estimated)18,765
13 Aug 2026 (estimated)18,765
6 Aug 2026 (estimated)18,294
30 Jul 2026 (estimated)18,271
23 Jul 2026 (estimated)18,271
16 Jul 2026 (estimated)18,260
9 Jul 2026 (estimated)18,248

Installs

5.6k

Tracking since . A chart appears once there are 7 days of data.

Installs via skills.sh

Similar skills