All tags
Catalogue tag

#crawling

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

22 records
Tagged “crawling”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
97Exceptionalhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
Packagist · Hex · crates.io +2
94Exceptionalhealth index
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
AGPL-3.0Aug 4, 2026 · metrics 2.10.0
NuGet
92Excellenthealth index
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,916Aug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm · PyPI
89Excellenthealth index
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,117Aug 24, 2026
AGPL-3.0Aug 24, 2026 · metrics 2.10.0
npm
86Excellenthealth index
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/moAug 5, 2026
AGPL-3.0Aug 5, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
linkchecker/linkchecker
check links in web documents or full websites
Python★ 1,067↓ 248.6K/moJul 30, 2026
GPL-2.0Jul 30, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
npm · crates.io · Packagist +1
78Goodhealth index
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
77Goodhealth index
tryAGI/Firecrawl
Generated C# SDK based on official Firecrawl OpenAPI specification
C#★ 4Sep 2, 2026
MITSep 2, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 177Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
57Moderatehealth index
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/moAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
Packagist · npm
50Moderatehealth index
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
Packagist
42Weakhealth index
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
Packagist
42Weakhealth index
firecrawl/firecrawl-php
No repository description published.
PHP★ 2↓ 2,586/moAug 19, 2026
No licenseAug 19, 2026 · metrics 2.10.0
Packagist
41Weakhealth index
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,078Jul 27, 2026
Custom licenseJul 27, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8KAug 12, 2026
BSD-3-ClauseAug 12, 2026 · metrics 2.10.0
npm
29At Riskhealth index
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 9Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
Packagist
28At Riskhealth index
roach-php/core
The complete web scraping toolkit for PHP.
PHP★ 1,455↓ 12K/moJul 27, 2026
No licenseJul 27, 2026 · metrics 2.10.0
npm
20At Riskhealth index
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0