Todas las etiquetas
Etiqueta del catálogo

#web-scraping

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

61 registros
Con la etiqueta «web-scraping»Ordenado por índice de salud
PyPI · npm
84Buenoíndice de salud
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
npm
84Buenoíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8K19 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0
PyPI · npm
82Buenoíndice de salud
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 17320 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
PyPI
81Buenoíndice de salud
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/mes14 jul 2026
BSD-3-Clause14 jul 2026 · métricas 1.13.0
81Buenoíndice de salud
mendableai/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 151.8K16 jul 2026
AGPL-3.016 jul 2026 · métricas 1.13.0
80Buenoíndice de salud
firecrawl/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript★ 149.9K↓ 0/mes13 jul 2026
AGPL-3.013 jul 2026 · métricas 1.13.0
PyPI
78Buenoíndice de salud
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6093↓ 36.1M/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
PyPI
75Buenoíndice de salud
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 12.9K↓ 3M/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
npm
73Buenoíndice de salud
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3201↓ 3522/mes22 jul 2026
Licencia propia22 jul 2026 · métricas 1.13.0
PyPI
72Buenoíndice de salud
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 631218 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
72Buenoíndice de salud
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2623/mes19 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0
PyPI · npm
72Buenoíndice de salud
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 1518 jul 2026
MIT18 jul 2026 · métricas 1.13.0
crates.io
71Buenoíndice de salud
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2601↓ 52.8K/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
Go
70Buenoíndice de salud
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 478115 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
70Buenoíndice de salud
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 20920 jul 2026
MIT20 jul 2026 · métricas 1.13.0
PyPI · npm
69Moderadoíndice de salud
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · TypeScript · C#★ 28.5K↓ 828.1K/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
npm
68Moderadoíndice de salud
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 182618 jul 2026
MIT18 jul 2026 · métricas 1.13.0
npm · crates.io
68Moderadoíndice de salud
olo-dot-io/uni-cli
Operations substrate for AI agents that use real software: 311 sites/tools, logged-in browsers, desktop apps, local tools, MCP, policy, evidence, AgentEnvelope v2, and self-repair.
TypeScript★ 87↓ 1161/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
npm
68Moderadoíndice de salud
vmoranv/jshookmcp
js hook toolkit that all you need
TypeScript★ 1756↓ 3417/mes18 jul 2026
AGPL-3.018 jul 2026 · métricas 1.13.0
Go · PyPI
68Moderadoíndice de salud
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 4220 jul 2026
MIT20 jul 2026 · métricas 1.13.0
Go
67Moderadoíndice de salud
HappyHackingSpace/dit
HTML page, form and field type classifier using ML (LogReg + CRF)
Go★ 1722 jul 2026
MIT22 jul 2026 · métricas 1.13.0
npm
67Moderadoíndice de salud
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4236/mes23 jul 2026
MIT23 jul 2026 · métricas 1.13.0
crates.io
66Moderadoíndice de salud
0x676e67/wreq-util
Common utilities for wreq
Rust★ 88↓ 221.4K/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI
65Moderadoíndice de salud
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 154↓ 13.4M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
npm
64Moderadoíndice de salud
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2900/mes17 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
crates.io
64Moderadoíndice de salud
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 9↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
crates.io · npm · Packagist +1
64Moderadoíndice de salud
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 140↓ 0/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
Go · npm
62Moderadoíndice de salud
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 38818 jul 2026
MIT18 jul 2026 · métricas 1.13.0
Go
62Moderadoíndice de salud
ChristopherDavenport/unblink
Pure-Go browser for AI: fetch pages, optionally render JS (no Chromium), get clean token-budgeted Markdown over MCP
Go★ 722 jul 2026
MIT22 jul 2026 · métricas 1.13.0
Go · npm
62Moderadoíndice de salud
felixgeelhaar/scout
Browser automation, one binary. The simpler alternative to Playwright — no Node, no Python, no runtime. Library, CLI, MCP server, and chat UI for any AI agent.
Go★ 615 jul 2026
MIT15 jul 2026 · métricas 1.13.0