Todas las etiquetas
Etiqueta del catálogo

#web-crawling

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

9 registros
Con la etiqueta «web-crawling»Ordenado por índice de salud
npm
98Excepcionalíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 942612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · npm
97Excepcionalíndice de salud
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 17320 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
npm
80Excelenteíndice de salud
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2614↓ 28.7K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
67Buenoíndice de salud
superagents-lab/search1api-mcp
Official Search1API MCP server for web search, news, crawling, sitemaps, and trends—hosted with OAuth 2.1 or local via npm.
TypeScript · JavaScript★ 173↓ 2722/mes24 ago 2026
MIT24 ago 2026 · métricas 2.10.0
npm · crates.io · Go +1
63Moderadoíndice de salud
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/mes16 jul 2026
MIT16 jul 2026 · métricas 2.10.0
npm
47Débilíndice de salud
NovadaLabs/novada-mcp
One MCP server for all web data — search, scrape, crawl, proxy, and AI research in a single npx install. Works with Claude, Cursor, and any MCP client.
TypeScript · HTML★ 2↓ 5804/mes16 jul 2026
Sin licencia16 jul 2026 · métricas 2.10.0