Web to Markdown
SoftaworksConverts webpage URLs to clean Markdown with a local web2md CLI that renders JavaScript-heavy pages.
Category · 66 skills
The skills in the directory filed under Data, in the editor's order.
Converts webpage URLs to clean Markdown with a local web2md CLI that renders JavaScript-heavy pages.
Scrapes data from Instagram, TikTok, YouTube, LinkedIn, Google Maps, Reddit and more than twenty other platforms through Apify Actors.
Plans and reviews MySQL and InnoDB schemas, indexing, query tuning, transactions and operations.
Explains Neki, PlanetScale's sharded Postgres product, for Postgres scaling and sharding work.
Covers PostgreSQL best practices: query optimisation, connection troubleshooting and performance.
Covers Vitess best practices: query optimisation, sharding, VSchema configuration and keyspace management.
Navigates websites on its own with Firecrawl and extracts structured data across pages.
Extracts many pages from one site or section in bulk with Firecrawl.
Saves a site or section as local markdown files and screenshots with Firecrawl, for offline reference.
Discovers and lists a site's URLs with Firecrawl, with search filtering.
Reads a known webpage with Firecrawl and returns its content or structured results.
Attaches a DuckDB database file, explores its schema and saves the session state for later queries.
Converts a data file between formats such as CSV, Parquet, JSON, Excel and GeoJSON with DuckDB.
Searches the DuckDB and DuckLake documentation and blog posts through a locally cached full-text index.
Runs SQL, or answers natural-language questions, against an attached DuckDB database or directly against files.
Reads and profiles a data file or remote URL in CSV, JSON, Parquet, Avro, Excel, spatial or SQLite format with DuckDB.
Explores and queries data on S3, Cloudflare R2, GCS, MinIO and other S3-compatible storage with DuckDB.
Runs analytical SQL on local files, URLs, S3 paths and remote databases with chDB, without setting up a server.
Helps design ClickHouse architectures and choose between ingestion and modelling patterns for a given workload.
Reviews ClickHouse schemas, queries and configurations against 31 documented rules.
Crawls a website and extracts content from many pages through the Tavily CLI.
Extracts clean markdown or text from specific URLs through the Tavily CLI.
Discovers and lists all the URLs on a website without extracting content, through the Tavily CLI.
Tunes MongoDB client connection settings such as pools and timeouts for any supported driver.