All tags
Catalogue tag

#web-crawler

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

21 records
Tagged “web-crawler”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
Packagist · Hex · crates.io +2
94Exceptionalhealth index
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
AGPL-3.0Aug 4, 2026 · metrics 2.10.0
npm
89Excellenthealth index
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7,332↓ 492.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm · PyPI
89Excellenthealth index
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,117Aug 24, 2026
AGPL-3.0Aug 24, 2026 · metrics 2.10.0
npm
87Excellenthealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm · PyPI · crates.io
87Excellenthealth index
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/moAug 3, 2026
AGPL-3.0Aug 3, 2026 · metrics 2.10.0
npm
80Excellenthealth index
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 1,826Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
Maven
78Goodhealth index
Norconex/crawler
Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
Java★ 204Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
Go
78Goodhealth index
kinorai/omnifeed
LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible
Go★ 6Sep 3, 2026
MITSep 3, 2026 · metrics 2.10.0
npm · crates.io · Packagist +1
78Goodhealth index
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
npm · crates.io
75Goodhealth index
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/moJul 26, 2026
AGPL-3.0Jul 26, 2026 · metrics 2.10.0
npm · PyPI · crates.io
75Goodhealth index
Goldziher/basemind
Full AI context and content layer for coding agents over one MCP server — tree-sitter code-map, document RAG, shared memory, multi-agent comms, web crawl, git history + blame. 300+ languages, 10+ agent harnesses, pure Rust.
Rust★ 69↓ 13.6K/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
Go
67Goodhealth index
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 207Aug 2, 2026
MITAug 2, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
npm · crates.io · Go +1
63Moderatehealth index
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
npm
41Weakhealth index
Aegis-Runner/AegisRunner
AegisRunner crawls any website from a single URL, discovers every page, form, and interaction, then uses AI to generate a complete Playwright test suite. Also runs accessibility, SEO, security, and performance audits on every page — no recording, no scripting, no setup.
JavaScript★ 0↓ 8,920/moJul 25, 2026
MITJul 25, 2026 · metrics 2.10.0
Maven
23At Riskhealth index
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4KAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0