All tags
Catalogue tag

#web-scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

105 records
Tagged “web-scraping”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
97Exceptionalhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
Packagist · PyPI
96Exceptionalhealth index
php-curl-class/php-curl-class
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
PHP★ 3,301↓ 140.4K/moJul 20, 2026
UnlicenseJul 20, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
Packagist · Hex · crates.io +2
94Exceptionalhealth index
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
AGPL-3.0Aug 4, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6,395↓ 45M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
93Exceptionalhealth index
browserbase/stagehand
The SDK For Browser Agents
TypeScript · MDX★ 23.7K↓ 4.8M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
npm · Go
92Excellenthealth index
pinchtab/pinchtab
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Go · Shell★ 10.1K↓ 7,120/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
Maven
90Excellenthealth index
jhy/jsoup
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Java · HTML★ 11.4KAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
89Excellenthealth index
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7,332↓ 492.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm · PyPI
88Excellenthealth index
MODSetter/SurfSense
Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9
Python · TypeScript★ 15.8KAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/moSep 4, 2026
MITSep 4, 2026 · metrics 2.10.0
npm
87Excellenthealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
npm
87Excellenthealth index
jo-inc/camofox-browser
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8,924↓ 243.1K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm · PyPI · crates.io
87Excellenthealth index
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/moAug 3, 2026
AGPL-3.0Aug 3, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2,623/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
npm
86Excellenthealth index
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/moAug 5, 2026
AGPL-3.0Aug 5, 2026 · metrics 2.10.0
Go
86Excellenthealth index
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 5,641Aug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm
86Excellenthealth index
kepano/defuddle
Get the main content of any page as Markdown.
TypeScript★ 9,191↓ 2.2M/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI · npm
86Excellenthealth index
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 15Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 209Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
PyPI · crates.io
81Excellenthealth index
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 1,423Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0