Optimize for GPU skill: what it does and how to install it
Optimises scientific Python workloads for NVIDIA GPUs with CUDA and RAPIDS, choosing suitable libraries, validating results and benchmarking end-to-end speed.
Summary generated from the skill's documentation.
Install
$ npx skills add K-Dense-AI/scientific-agent-skills --skill optimize-for-gpuRun it in a terminal. If your agent is already running, start a new session so it picks the skill up.
About this skill
What it does. Profiles a CPU baseline, selects the least disruptive CUDA or RAPIDS path, moves data through a coherent GPU pipeline, and compares outputs with numerical tolerances. It benchmarks synchronously, including transfers and memory, then keeps, revises or rejects the port.
When to use it. For large scientific Python workloads involving arrays, dataframes, machine learning, graphs, images, simulations, vector search or GPU file I/O. It keeps a CPU path for small, sequential or unsupported workloads and does not port code already well served by a GPU-native framework.
History
Repo stars
47.3kAbout +18.6k since 9 Jul 2026
Before 1 Oct 2026 the curve is estimated from public event data.
Stars are counted for the whole repository, which holds 50 skills.
Show as a table
| Date | Repo stars |
|---|---|
| 1 Oct 2026 | 47,271 |
| 24 Sept 2026 (estimated) | 46,976 |
| 17 Sept 2026 (estimated) | 46,257 |
| 10 Sept 2026 (estimated) | 41,836 |
| 3 Sept 2026 (estimated) | 30,545 |
| 27 Aug 2026 (estimated) | 29,347 |
| 20 Aug 2026 (estimated) | 29,160 |
| 13 Aug 2026 (estimated) | 29,027 |
| 6 Aug 2026 (estimated) | 28,947 |
| 30 Jul 2026 (estimated) | 28,894 |
| 23 Jul 2026 (estimated) | 28,894 |
| 16 Jul 2026 (estimated) | 28,841 |
| 9 Jul 2026 (estimated) | 28,708 |
Installs
1.7k
Tracking since . A chart appears once there are 7 days of data.
Installs via skills.sh
Similar skills
- Vercel OptimizeFinds cost and performance improvements for deployed Vercel projects from metrics, usage and code scans.
- Hugging Face GradioBuilds Gradio web UIs and demos in Python: components, event listeners, layouts and chatbots.
- Hugging Face Local ModelsHelps select and run models locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA or ROCm.