All tags
Catalogue tag

#crawling

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

9 records
Tagged “crawling”Ranked by health index
PyPI · npm
84Goodhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
84Goodhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8KJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI
81Goodhealth index
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/moJul 14, 2026
BSD-3-ClauseJul 14, 2026 · metrics 1.13.0
76Goodhealth index
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,900Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · npm
68Moderatehealth index
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/moJul 19, 2026
MITJul 19, 2026 · metrics 1.13.0
crates.io · npm · Packagist +1
64Moderatehealth index
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 140↓ 0/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 177Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
Packagist
46At riskhealth index
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
npm
35At riskhealth index
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 9Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0