All tags
Catalogue tag

#web-scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

107 records
Tagged “web-scraping”Ranked by health index
PyPI
81Excellenthealth index
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
PyPI · crates.io · npm
81Excellenthealth index
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
npm · crates.io
81Excellenthealth index
olo-dot-io/uni-cli
Operations substrate for AI agents that use real software: 311 sites/tools, logged-in browsers, desktop apps, local tools, MCP, policy, evidence, AgentEnvelope v2, and self-repair.
TypeScript★ 87↓ 1,161/moJul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go
80Excellenthealth index
HappyHackingSpace/dit
HTML page, form and field type classifier using ML (LogReg + CRF)
Go★ 17Jul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
npm
80Excellenthealth index
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
Go
80Excellenthealth index
karust/openserp
Self-hosted SERP API for AI, SEO & automation. Browser-rendered Google, Bing, Yandex, Baidu, DuckDuckGo and Ecosia search with page extraction 🎉
Go★ 1,262Aug 14, 2026
MITAug 14, 2026 · metrics 2.10.0
Go · npm
80Excellenthealth index
klarlabs-studio/scout
Browser automation, one binary. The simpler alternative to Playwright — no Node, no Python, no runtime. Library, CLI, MCP server, and chat UI for any AI agent.
Go★ 11Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
npm
80Excellenthealth index
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 1,826Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
Go · PyPI
80Excellenthealth index
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 42Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
crates.io
78Goodhealth index
0x676e67/wreq-util
Common utilities for wreq
Rust★ 90↓ 209.7K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
npm
78Goodhealth index
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 155↓ 15M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
Go
78Goodhealth index
kinorai/omnifeed
LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible
Go★ 6Sep 3, 2026
MITSep 3, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
rushter/selectolax
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 1,656Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
npm
78Goodhealth index
vmoranv/jshookmcp
js hook toolkit that all you need
TypeScript★ 1,756↓ 3,417/moJul 18, 2026
AGPL-3.0Jul 18, 2026 · metrics 2.10.0
npm · crates.io · Packagist +1
78Goodhealth index
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
npm
77Goodhealth index
n1byn1kt/apitap
CLI, MCP server, and npm library that turns any website into an API — no docs, no SDK, no browser.
TypeScript★ 123↓ 2,030/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
npm · crates.io
75Goodhealth index
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/moJul 26, 2026
AGPL-3.0Jul 26, 2026 · metrics 2.10.0
npm
73Goodhealth index
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/moJul 26, 2026
No licenseJul 26, 2026 · metrics 2.10.0
npm
71Goodhealth index
F4RAN/cuimp-ts
A Node.js wrapper for curl-impersonate that allows you to make HTTP requests that mimic real browser behavior, bypassing many anti-bot protections.
TypeScript★ 89↓ 22.7K/moAug 1, 2026
No licenseAug 1, 2026 · metrics 2.10.0
npm
71Goodhealth index
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2,900/moJul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
npm
71Goodhealth index
intoli/user-agents
A JavaScript library for generating random user agents with data that's updated daily.
TypeScript · JavaScript★ 1,189↓ 812.5K/moSep 5, 2026
Custom licenseSep 5, 2026 · metrics 2.10.0
npm
71Goodhealth index
scrape-badger/scrapebadger-node
Official Node.js SDK for ScrapeBadger - Async web scraping APIs for Twitter and more
TypeScript★ 2↓ 3,672/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
Go · npm
69Goodhealth index
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 388Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
Go
69Goodhealth index
ChristopherDavenport/unblink
Pure-Go browser for AI: fetch pages, optionally render JS (no Chromium), get clean token-budgeted Markdown over MCP
Go★ 7Jul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
ebarti/jobstreaming
Concurrent, resumable job collection for Python with typed events, durable checkpoints, normalized models, and a DataFrame API.
Python★ 0↓ 2,749/moAug 29, 2026
MITAug 29, 2026 · metrics 2.10.0
Go
67Goodhealth index
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 207Aug 2, 2026
MITAug 2, 2026 · metrics 2.10.0
Go · npm
67Goodhealth index
sardanioss/httpcloak
Go HTTP client with browser-identical TLS/HTTP2 fingerprinting. Bypass bot detection by perfectly mimicking Chrome, Firefox, and Safari at the cryptographic level (JA3/JA4, Akamai fingerprint, header order). Supports HTTP/1.1, HTTP/2, HTTP/3, sessions, cookies, and proxies.
Go · C#★ 1,161Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0