Todas las etiquetas
Etiqueta del catálogo

#web-scraping

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

105 registros
Con la etiqueta «web-scraping»Ordenado por índice de salud
npm
98Excepcionalíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 942612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · npm
97Excepcionalíndice de salud
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
Packagist · PyPI
96Excepcionalíndice de salud
php-curl-class/php-curl-class
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
PHP★ 3301↓ 140.4K/mes20 jul 2026
Unlicense20 jul 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K4 ago 2026
BSD-3-Clause4 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 17320 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
Packagist · Hex · crates.io +2
94Excepcionalíndice de salud
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3278/mes4 ago 2026
AGPL-3.04 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6395↓ 45M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
npm
93Excepcionalíndice de salud
browserbase/stagehand
The SDK For Browser Agents
TypeScript · MDX★ 23.7K↓ 4.8M/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
npm · Go
92Excelenteíndice de salud
pinchtab/pinchtab
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Go · Shell★ 10.1K↓ 7120/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6718↓ 14M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
Maven
90Excelenteíndice de salud
jhy/jsoup
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Java · HTML★ 11.4K27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
npm
89Excelenteíndice de salud
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7332↓ 492.7K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm · PyPI
88Excelenteíndice de salud
MODSetter/SurfSense
Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9
Python · TypeScript★ 15.8K5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI
88Excelenteíndice de salud
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1974↓ 15.7K/mes4 sept 2026
MIT4 sept 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3201↓ 3522/mes22 jul 2026
Licencia propia22 jul 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
jo-inc/camofox-browser
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8924↓ 243.1K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
crates.io
87Excelenteíndice de salud
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2669↓ 42K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm · PyPI · crates.io
87Excelenteíndice de salud
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9070/mes3 ago 2026
AGPL-3.03 ago 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2623/mes19 jul 2026
Apache-2.019 jul 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/mes5 ago 2026
AGPL-3.05 ago 2026 · métricas 2.10.0
Go
86Excelenteíndice de salud
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 564128 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
kepano/defuddle
Get the main content of any page as Markdown.
TypeScript★ 9191↓ 2.2M/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI · npm
86Excelenteíndice de salud
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 1518 jul 2026
MIT18 jul 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 20920 jul 2026
MIT20 jul 2026 · métricas 2.10.0
PyPI · npm
83Excelenteíndice de salud
sportsdataverse/sportsdataverse-py
sportsdataverse python package
Python★ 111↓ 15.8K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
PyPI · crates.io
81Excelenteíndice de salud
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 142313 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/mes27 jul 2026
MIT27 jul 2026 · métricas 2.10.0