npm80Excellenthealth index
TypeScript★ 22↓ 3,745/moJul 20, 2026
MrAdex77/google-play-scraperGoogle Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/moJul 23, 2026

Norconex/crawlerNorconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
Java★ 204Sep 5, 2026
rushter/selectolaxPython binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 1,656Jul 20, 2026
C++ · JavaScript★ 2↓ 254.8K/moJul 16, 2026

JaCraig/SpideyA multi threaded web crawler library that is generic enough to allow different engines to be swapped in.
C#★ 16Aug 21, 2026
Packagist77Goodhealth index
spatie/robots-txtDetermine if a page may be crawled from robots.txt, robots meta tags and robot headers
PHP★ 258↓ 785.7K/moAug 4, 2026
npm · crates.io75Goodhealth index
0xMassi/webclawFast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/moJul 26, 2026
C#★ 0Aug 31, 2026
Liu233w/ojhunt-liteA lightweight async Python tool for querying Online Judge (OJ) statistics across multiple platforms. Track your accepted problems (AC) and total submissions from 29+ competitive programming platforms.
Python★ 5↓ 3,655/moJul 29, 2026
Python★ 726↓ 2,592/moJul 19, 2026
firecrawl/cliCLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/moJul 26, 2026
iannuttall/seoThe only SEO skill your agent needs. 70+ SEO audit tools through a local CLI and MCP server, using your own crawl, Search Console, and GA4 data.
TypeScript · Astro★ 81↓ 4,913/moAug 2, 2026
TypeScript · HTML · CSS★ 23↓ 14.6K/moSep 5, 2026
Go · Python★ 52Jul 15, 2026
adbar/courlanClean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 177Jul 21, 2026
Java★ 808Jul 17, 2026
Go★ 207Aug 2, 2026

krau/ManyACGCollect, Download, Organize and Share your Favorite Anime Artworks.
Go★ 115Aug 22, 2026
TypeScript★ 1↓ 4,014/moJul 26, 2026
npm63Moderatehealth index
ODATANO/NIGHTGATEODATANO-NIGHTGATE connects SAP systems to the Midnight privacy chain via a CAP-based OData V4 API, enabling confidential on-chain data access and privacy-preserving transaction execution
TypeScript · JavaScript★ 3↓ 3,520/moJul 18, 2026
npm63Moderatehealth index
TypeScript · JavaScript★ 11↓ 1,369/moJul 15, 2026
RubyGems63Moderatehealth index
Ruby★ 148Sep 6, 2026
npm63Moderatehealth index

mysleekdesigns/crawlforge-mcp28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/moSep 5, 2026
npm · crates.io · Go +163Moderatehealth index
Rust · Python · Go★ 26↓ 16.3K/moJul 16, 2026
npm63Moderatehealth index
thecodrr/fdir⚡ The fastest directory crawler & globbing library for NodeJS. Crawls 1m files in < 1s
TypeScript★ 1,728↓ 630M/moAug 4, 2026
Go★ 1Jul 21, 2026
NuGet62Moderatehealth index

zzzprojects/html-agility-packHtml Agility Pack (HAP) is a free and open-source HTML parser written in C# to read/write DOM and supports plain XPATH or XSLT. It is a .NET code library that allows you to parse "out of the web" HTML files.
C# · HTML★ 2,846Aug 22, 2026
PyPI60Moderatehealth index
moskrc/crawlerdetect🕷CrawlerDetect is a Python library designed to identify bots, crawlers, and spiders by analyzing their user agents.
Python★ 44Jul 29, 2026
PyPI59Moderatehealth index
Python · Shell★ 299Jul 21, 2026