All tags
Catalogue tag

#crawler

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

86 records
Tagged “crawler”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
Packagist · Hex · crates.io +2
94Exceptionalhealth index
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
AGPL-3.0Aug 4, 2026 · metrics 2.10.0
NuGet
92Excellenthealth index
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,916Aug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
npm
91Excellenthealth index
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
PyPI · npm
90Excellenthealth index
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2,566↓ 2.2M/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
mikf/gallery-dl
Command-line program to download image galleries and collections from several image hosting sites
Python★ 19.1KAug 5, 2026
GPL-2.0Aug 5, 2026 · metrics 2.10.0
Maven
89Excellenthealth index
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
Java★ 1,936Aug 9, 2026
GPL-3.0Aug 9, 2026 · metrics 2.10.0
Packagist
89Excellenthealth index
tomasnorre/crawler
Libraries and scripts for crawling the TYPO3 page tree. Used for re-caching, re-indexing, publishing applications etc.
PHP★ 57↓ 8,254/moAug 22, 2026
GPL-3.0Aug 22, 2026 · metrics 2.10.0
npm · PyPI
89Excellenthealth index
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,117Aug 24, 2026
AGPL-3.0Aug 24, 2026 · metrics 2.10.0
Packagist · npm
88Excellenthealth index
eliashaeussler/cache-warmup
🔥 PHP library to warm up caches of URLs located in XML sitemaps
PHP★ 77↓ 14.1K/moJul 16, 2026
GPL-3.0Jul 16, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/moSep 4, 2026
MITSep 4, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
0x676e67/wreq
An ergonomic, privacy-aware Rust HTTP Client
Rust★ 1,004↓ 220.3K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
87Excellenthealth index
DedSecInside/TorBot
Dark Web OSINT Tool
Python★ 4,726Aug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
npm
87Excellenthealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm · PyPI · crates.io
87Excellenthealth index
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/moAug 3, 2026
AGPL-3.0Aug 3, 2026 · metrics 2.10.0
npm
87Excellenthealth index
zorillajs/zorilla
Zorilla is a modular plugin framework that extends Puppeteer and Playwright with additional functionality through a clean plugin architecture. Build powerful browser automation with composable plugins for stealth mode, captcha solving, ad blocking, and much more.
TypeScript · JavaScript★ 27↓ 15.3K/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
npm
86Excellenthealth index
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/moAug 5, 2026
AGPL-3.0Aug 5, 2026 · metrics 2.10.0
Go
86Excellenthealth index
openclaw/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go · Python★ 56Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
NuGet
84Excellenthealth index
51Degrees/device-detection-dotnet
Device detection services for 51Degrees Pipeline
C# · C++★ 11Aug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
npm
83Excellenthealth index
microlinkhq/top-user-agents
An always up-to-date list of the top 100 HTTP user-agents most used over the Internet.
JavaScript★ 357↓ 20.3K/moAug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
Packagist
83Excellenthealth index
spatie/crawler
https://spatie.be/docs/crawler
PHP★ 2,829↓ 751.2K/moJul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
PyPI · crates.io
81Excellenthealth index
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 1,423Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
z-mio/ParseHub
轻量、异步、开箱即用的社交媒体聚合解析库
Python★ 152↓ 2,797/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
Packagist
80Excellenthealth index
JayBizzle/Crawler-Detect
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
PHP★ 2,401↓ 2.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0