All tags
Catalogue tag

#scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

43 records
Tagged “scraping”Ranked by health index
PyPI · npm
84Goodhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
84Goodhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8KJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI · npm
82Goodhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
PyPI
81Goodhealth index
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/moJul 14, 2026
BSD-3-ClauseJul 14, 2026 · metrics 1.13.0
81Goodhealth index
mendableai/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 151.8KJul 16, 2026
AGPL-3.0Jul 16, 2026 · metrics 1.13.0
80Goodhealth index
firecrawl/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript★ 149.9K↓ 0/moJul 13, 2026
AGPL-3.0Jul 13, 2026 · metrics 1.13.0
crates.io
79Goodhealth index
plabayo/rama
modular service framework to move and transform network packets
Rust★ 1,088↓ 642.9K/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
npm
73Goodhealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 1.13.0
PyPI
72Goodhealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,312Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
crates.io
71Goodhealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,601↓ 52.8K/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
npm
70Goodhealth index
achiya-automation/safari-mcp
Native Safari browser automation for AI agents. 80 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chrome DevTools MCP with 40-60% less CPU/heat on Apple Silicon.
JavaScript★ 151↓ 6,053/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI
70Goodhealth index
blisspixel/primr
Turn any company URL into a strategic intelligence brief. Adaptive scraping + AI-powered research and synthesis.
Python★ 3↓ 7,380/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI
70Goodhealth index
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,038↓ 0/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
PyPI · npm
69Moderatehealth index
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · TypeScript · C#★ 28.5K↓ 828.1K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · npm
68Moderatehealth index
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/moJul 19, 2026
MITJul 19, 2026 · metrics 1.13.0
Packagist · npm
68Moderatehealth index
playwright-php/playwright
Playwright PHP library for browser automation: navigation, E2E tests, assertions, screenshots, and so much more!
PHP★ 204↓ 12.7K/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
crates.io
64Moderatehealth index
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 9↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
Go · npm
62Moderatehealth index
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 388Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
Packagist
62Moderatehealth index
php-embed/Embed
Get info from any web service or page
PHP★ 2,142↓ 290.6K/moJul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
npm
61Moderatehealth index
scrape-badger/scrapebadger-node
Official Node.js SDK for ScrapeBadger - Async web scraping APIs for Twitter and more
TypeScript★ 2↓ 2,131/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
npm
61Moderatehealth index
stacksjs/ts-web-scraper
A powerful, type-safe web scraping library for TypeScript.
TypeScript★ 11↓ 177.5K/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
npm · RubyGems
60Moderatehealth index
LeonTing1010/tap
Capture a logged-in browser task once — replay it forever at zero LLM tokens. Local-first browser-automation MCP for Claude Code, Cursor & any MCP host; credentials never leave your machine.
JavaScript · TypeScript★ 12Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm
60Moderatehealth index
Lincoln504/pi-research
Web research for pi with smart and safe tooling + agent system
TypeScript★ 20↓ 2,031/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
Go · npm
59Moderatehealth index
alpkeskin/rota
A high-performance proxy rotation engine with automated IP management and real-time health monitoring
Go · TypeScript★ 508Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
Go
59Moderatehealth index
anatolykoptev/go-browser
Stealth Chrome runtime for autonomous AI agents — anti-detection browser for authorized automation tasks (bookings, form filling, research)
Go · JavaScript★ 2Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
58Moderatehealth index
clearcotelabs/clearcote-browser
Open-source stealth Chromium 149 with engine-level fingerprint spoofing - de-Googled, drop-in Playwright, fully buildable and verifiable from source.
Python · TypeScript · C#★ 31↓ 141/moJul 18, 2026
BSD-3-ClauseJul 18, 2026 · metrics 1.13.0
Go
58Moderatehealth index
wasylq/fss
fss or Full Studio Scraper - scrape all videos metadata of your favourite performers or studios
Go★ 5↓ 0/moJul 13, 2026
GPL-3.0Jul 13, 2026 · metrics 1.13.0
PyPI
56Moderatehealth index
MathiasPaulenko/wavexis
Browser automation CLI — wraps cdpwave and bidiwave, no Node.js, no Chromium download. CDP & WebDriver BiDi backends, 80+ commands, async core, serve mode, session recording.
Python★ 1↓ 6,037/moJul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
Go
51Moderatehealth index
Hyper-Solutions/hyper-sdk-go
Go SDK for Bot Protection Bypass - Automate Akamai, Incapsula, Kasada, and DataDome. No browsers required. Solve challenges and generate valid sensors/cookies via API.
Go★ 62Jul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
npm
51Moderatehealth index
Hyper-Solutions/hyper-sdk-js
JavaScript / TypeScript SDK for Bot Protection Bypass - Automate Akamai, Incapsula, Kasada, and DataDome. No browsers required. Solve challenges and generate valid sensors/cookies via API.
TypeScript★ 54↓ 31.6K/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0