跳过主要内容

Category · 66 skills

Data skills

The skills in the directory filed under Data, in the editor's order.

Data skills

Defuddle

Steph Ango

Extracts clean Markdown from HTML pages with the Defuddle CLI.

DataGitHub stars: 49.1k

Biopython

K-Dense

Molecular biology toolkit: sequence manipulation, FASTA, GenBank and PDB parsing, phylogenetics and NCBI access.

DataGitHub stars: 47.3k

Runs exploratory analysis on scientific data files: profiles, missing-data audits and outlier checks.

DataGitHub stars: 47.3k

Matplotlib

K-Dense

Reference for Matplotlib: fine-grained control over plot elements and export to PNG, PDF or SVG for publication.

DataGitHub stars: 47.3k

NetworkX

K-Dense

Creates, analyses and visualises networks and graphs in Python with NetworkX.

DataGitHub stars: 47.3k

Polars

K-Dense

Reference for Polars, the Python DataFrame library: expressions, lazy queries, streaming and migration from pandas.

DataGitHub stars: 47.3k

PyMC

K-Dense

Builds Bayesian models with PyMC: hierarchical models, MCMC, variational inference and model comparison.

DataGitHub stars: 47.3k

RDKit

K-Dense

Cheminformatics with RDKit: SMILES parsing, molecular descriptors, fingerprints, substructure search and reactions.

DataGitHub stars: 47.3k

Scanpy

K-Dense

Runs the standard single-cell RNA-seq pipeline with Scanpy: QC, normalisation, clustering and differential expression.

DataGitHub stars: 47.3k

Creates and audits publication-ready scientific figures with Matplotlib, Seaborn or Plotly.

DataGitHub stars: 47.3k

Machine learning in Python with scikit-learn: classification, regression, clustering, model evaluation and tuning.

DataGitHub stars: 47.3k

Seaborn

K-Dense

Statistical visualisation with Seaborn: distributions, relationships and categorical comparisons.

DataGitHub stars: 47.3k

Guides statistical analysis of research data: test selection, assumption checks, effect sizes and APA-formatted reporting.

DataGitHub stars: 47.3k

Statistical modelling with statsmodels: OLS, GLM, mixed models and ARIMA, with diagnostics and inference.

DataGitHub stars: 47.3k

Jupyter Notebook

OpenAIOfficial

Creates and edits Jupyter notebooks for experiments, explorations and tutorials from bundled templates.

DataGitHub stars: 27.8k

Baoyu URL to Markdown

Jim Liu 宝玉

Fetches any URL and converts it to markdown through Chrome, with adapters for X, YouTube transcripts and Hacker News.

DataGitHub stars: 26.2k

BigQuery Basics

GoogleOfficial

Manages BigQuery datasets, tables and jobs, and runs SQL queries for basic ingestion and analysis.

DataGitHub stars: 20.6k

Cloud SQL Basics

GoogleOfficial

Creates and explains Cloud SQL instances and databases for MySQL, PostgreSQL and SQL Server.

DataGitHub stars: 20.6k

Hugging Face Community Evals

Hugging FaceOfficial

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai and lighteval.

DataGitHub stars: 11.1k

Hugging Face Datasets

Hugging FaceOfficial

Works with the Hugging Face Dataset Viewer API: split metadata, row pagination, text search, filters and parquet downloads.

DataGitHub stars: 11.1k

Hugging Face Trackio

Hugging FaceOfficial

Tracks and visualises ML training experiments with Trackio: logged metrics, alerts and a live dashboard.

DataGitHub stars: 11.1k

Postgres best practices maintained by Supabase, for schema changes, queries, indexes and security on any Postgres.

DataGitHub stars: 2.7k

Designs SQL and NoSQL database schemas: normalisation, indexing, constraints and migration patterns.

DataGitHub stars: 2.5k

Other categories