Todas las etiquetas
Etiqueta del catálogo

#web-scraping

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

107 registros
Con la etiqueta «web-scraping»Ordenado por índice de salud
PyPI
81Excelenteíndice de salud
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/mes27 jul 2026
MIT27 jul 2026 · métricas 2.10.0
PyPI · crates.io · npm
81Excelenteíndice de salud
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6257/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
npm · crates.io
81Excelenteíndice de salud
olo-dot-io/uni-cli
Operations substrate for AI agents that use real software: 311 sites/tools, logged-in browsers, desktop apps, local tools, MCP, policy, evidence, AgentEnvelope v2, and self-repair.
TypeScript★ 87↓ 1161/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI · npm
80Excelenteíndice de salud
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go
80Excelenteíndice de salud
HappyHackingSpace/dit
HTML page, form and field type classifier using ML (LogReg + CRF)
Go★ 1722 jul 2026
MIT22 jul 2026 · métricas 2.10.0
npm
80Excelenteíndice de salud
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2614↓ 28.7K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
Go
80Excelenteíndice de salud
karust/openserp
Self-hosted SERP API for AI, SEO & automation. Browser-rendered Google, Bing, Yandex, Baidu, DuckDuckGo and Ecosia search with page extraction 🎉
Go★ 126214 ago 2026
MIT14 ago 2026 · métricas 2.10.0
Go · npm
80Excelenteíndice de salud
klarlabs-studio/scout
Browser automation, one binary. The simpler alternative to Playwright — no Node, no Python, no runtime. Library, CLI, MCP server, and chat UI for any AI agent.
Go★ 115 sept 2026
MIT5 sept 2026 · métricas 2.10.0
npm
80Excelenteíndice de salud
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 182618 jul 2026
MIT18 jul 2026 · métricas 2.10.0
Go · PyPI
80Excelenteíndice de salud
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 4220 jul 2026
MIT20 jul 2026 · métricas 2.10.0
crates.io
78Buenoíndice de salud
0x676e67/wreq-util
Common utilities for wreq
Rust★ 90↓ 209.7K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4236/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 155↓ 15M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
Go
78Buenoíndice de salud
kinorai/omnifeed
LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible
Go★ 63 sept 2026
MIT3 sept 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
rushter/selectolax
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 165620 jul 2026
MIT20 jul 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
vmoranv/jshookmcp
js hook toolkit that all you need
TypeScript★ 1756↓ 3417/mes18 jul 2026
AGPL-3.018 jul 2026 · métricas 2.10.0
npm · crates.io · Packagist +1
78Buenoíndice de salud
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
npm
77Buenoíndice de salud
n1byn1kt/apitap
CLI, MCP server, and npm library that turns any website into an API — no docs, no SDK, no browser.
TypeScript★ 123↓ 2030/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm · crates.io
75Buenoíndice de salud
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2066↓ 424/mes26 jul 2026
AGPL-3.026 jul 2026 · métricas 2.10.0
npm
73Buenoíndice de salud
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/mes26 jul 2026
Sin licencia26 jul 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
F4RAN/cuimp-ts
A Node.js wrapper for curl-impersonate that allows you to make HTTP requests that mimic real browser behavior, bypassing many anti-bot protections.
TypeScript★ 89↓ 22.7K/mes1 ago 2026
Sin licencia1 ago 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2900/mes17 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
intoli/user-agents
A JavaScript library for generating random user agents with data that's updated daily.
TypeScript · JavaScript★ 1189↓ 812.5K/mes5 sept 2026
Licencia propia5 sept 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
scrape-badger/scrapebadger-node
Official Node.js SDK for ScrapeBadger - Async web scraping APIs for Twitter and more
TypeScript★ 2↓ 3672/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
Go · npm
69Buenoíndice de salud
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 38818 jul 2026
MIT18 jul 2026 · métricas 2.10.0
Go
69Buenoíndice de salud
ChristopherDavenport/unblink
Pure-Go browser for AI: fetch pages, optionally render JS (no Chromium), get clean token-budgeted Markdown over MCP
Go★ 722 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
ebarti/jobstreaming
Concurrent, resumable job collection for Python with typed events, durable checkpoints, normalized models, and a DataFrame API.
Python★ 0↓ 2749/mes29 ago 2026
MIT29 ago 2026 · métricas 2.10.0
Go
67Buenoíndice de salud
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 2072 ago 2026
MIT2 ago 2026 · métricas 2.10.0
Go · npm
67Buenoíndice de salud
sardanioss/httpcloak
Go HTTP client with browser-identical TLS/HTTP2 fingerprinting. Bypass bot detection by perfectly mimicking Chrome, Firefox, and Safari at the cryptographic level (JA3/JA4, Akamai fingerprint, header order). Supports HTTP/1.1, HTTP/2, HTTP/3, sessions, cookies, and proxies.
Go · C#★ 116115 jul 2026
MIT15 jul 2026 · métricas 2.10.0