All tags
Catalogue tag

#web-scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

60 records
Tagged “web-scraping”Ranked by health index
PyPI · npm
84Goodhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
84Goodhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8KJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI · npm
82Goodhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
PyPI
81Goodhealth index
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/moJul 14, 2026
BSD-3-ClauseJul 14, 2026 · metrics 1.13.0
81Goodhealth index
mendableai/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 151.8KJul 16, 2026
AGPL-3.0Jul 16, 2026 · metrics 1.13.0
80Goodhealth index
firecrawl/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript★ 149.9K↓ 0/moJul 13, 2026
AGPL-3.0Jul 13, 2026 · metrics 1.13.0
PyPI
78Goodhealth index
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6,093↓ 36.1M/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
PyPI
75Goodhealth index
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 12.9K↓ 3M/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm
73Goodhealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 1.13.0
PyPI
72Goodhealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,312Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
72Goodhealth index
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2,623/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI · npm
72Goodhealth index
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 15Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
crates.io
71Goodhealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,601↓ 52.8K/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
Go
70Goodhealth index
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 4,781Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
PyPI
70Goodhealth index
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 209Jul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
PyPI · npm
69Moderatehealth index
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · TypeScript · C#★ 28.5K↓ 828.1K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm
68Moderatehealth index
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 1,826Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm · crates.io
68Moderatehealth index
olo-dot-io/uni-cli
Operations substrate for AI agents that use real software: 311 sites/tools, logged-in browsers, desktop apps, local tools, MCP, policy, evidence, AgentEnvelope v2, and self-repair.
TypeScript★ 87↓ 1,161/moJul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
npm
68Moderatehealth index
vmoranv/jshookmcp
js hook toolkit that all you need
TypeScript★ 1,756↓ 3,417/moJul 18, 2026
AGPL-3.0Jul 18, 2026 · metrics 1.13.0
Go · PyPI
68Moderatehealth index
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 42Jul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
Go
67Moderatehealth index
HappyHackingSpace/dit
HTML page, form and field type classifier using ML (LogReg + CRF)
Go★ 17Jul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
npm
67Moderatehealth index
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/moJul 23, 2026
MITJul 23, 2026 · metrics 1.13.0
crates.io
66Moderatehealth index
0x676e67/wreq-util
Common utilities for wreq
Rust★ 88↓ 221.4K/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
PyPI
65Moderatehealth index
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 154↓ 13.4M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
64Moderatehealth index
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2,900/moJul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
crates.io
64Moderatehealth index
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 9↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
crates.io · npm · Packagist +1
64Moderatehealth index
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 140↓ 0/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
Go · npm
62Moderatehealth index
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 388Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
Go
62Moderatehealth index
ChristopherDavenport/unblink
Pure-Go browser for AI: fetch pages, optionally render JS (no Chromium), get clean token-budgeted Markdown over MCP
Go★ 7Jul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
Go · npm
62Moderatehealth index
felixgeelhaar/scout
Browser automation, one binary. The simpler alternative to Playwright — no Node, no Python, no runtime. Library, CLI, MCP server, and chat UI for any AI agent.
Go★ 6Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0