All tags
Catalogue tag

#scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

82 records
Tagged “scraping”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
97Exceptionalhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
Packagist · Hex · crates.io +2
94Exceptionalhealth index
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
AGPL-3.0Aug 4, 2026 · metrics 2.10.0
crates.io
94Exceptionalhealth index
plabayo/rama
modular service framework to move and transform network packets
Rust★ 1,190↓ 730.4K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
freelawproject/juriscraper
An API to scrape American court websites for metadata.
HTML★ 623↓ 25K/moAug 18, 2026
BSD-2-ClauseAug 18, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI · npm
90Excellenthealth index
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2,566↓ 2.2M/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
browser-use/browser-harness
Browser Harness | Self-healing harness that enables LLMs to complete any task.
Python · HTML★ 16.5KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
soxoj/maigret
🕵️‍♂️ Collect a dossier on a person by username from 3000+ sites
Python · HTML★ 36.1K↓ 99.4K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,067↓ 111.2K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/moSep 4, 2026
MITSep 4, 2026 · metrics 2.10.0
npm
87Excellenthealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
PyPI · npm
87Excellenthealth index
daijro/camoufox
🦊 Anti-detect browser
C++ · Python · JavaScript★ 11.7K↓ 851.4K/moSep 6, 2026
MPL-2.0Sep 6, 2026 · metrics 2.10.0
npm
87Excellenthealth index
jo-inc/camofox-browser
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8,924↓ 243.1K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
86Excellenthealth index
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/moAug 5, 2026
AGPL-3.0Aug 5, 2026 · metrics 2.10.0
npm
83Excellenthealth index
achiya-automation/safari-mcp
Native Safari browser automation for AI agents. 80 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chrome DevTools MCP with 40-60% less CPU/heat on Apple Silicon.
JavaScript★ 151↓ 6,053/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
blisspixel/primr
Turn any company URL into a strategic intelligence brief. Adaptive scraping + AI-powered research and synthesis.
Python★ 3↓ 7,380/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
npm
83Excellenthealth index
microlinkhq/top-user-agents
An always up-to-date list of the top 100 HTTP user-agents most used over the Internet.
JavaScript★ 357↓ 20.3K/moAug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
PyPI · crates.io · npm
81Excellenthealth index
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
npm
80Excellenthealth index
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
gawel/pyquery
A jquery-like library for python
Python★ 2,381↓ 1.9M/moAug 3, 2026
Custom licenseAug 3, 2026 · metrics 2.10.0
Packagist · npm
78Goodhealth index
playwright-php/playwright
Playwright PHP library for browser automation: navigation, E2E tests, assertions, screenshots, and so much more!
PHP★ 204↓ 12.7K/moJul 18, 2026
MITJul 18, 2026 · metrics 2.10.0