Todas las etiquetas
Etiqueta del catálogo

#crawler

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

86 registros
Con la etiqueta «crawler»Ordenado por índice de salud
npm
98Excepcionalíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 942612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K4 ago 2026
BSD-3-Clause4 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 17320 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
Packagist · Hex · crates.io +2
94Excepcionalíndice de salud
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3278/mes4 ago 2026
AGPL-3.04 ago 2026 · métricas 2.10.0
NuGet
92Excelenteíndice de salud
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 391628 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6718↓ 14M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm
91Excelenteíndice de salud
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
PyPI · npm
90Excelenteíndice de salud
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2566↓ 2.2M/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
mikf/gallery-dl
Command-line program to download image galleries and collections from several image hosting sites
Python★ 19.1K5 ago 2026
GPL-2.05 ago 2026 · métricas 2.10.0
Maven
89Excelenteíndice de salud
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
Java★ 19369 ago 2026
GPL-3.09 ago 2026 · métricas 2.10.0
Packagist
89Excelenteíndice de salud
tomasnorre/crawler
Libraries and scripts for crawling the TYPO3 page tree. Used for re-caching, re-indexing, publishing applications etc.
PHP★ 57↓ 8254/mes22 ago 2026
GPL-3.022 ago 2026 · métricas 2.10.0
npm · PyPI
89Excelenteíndice de salud
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 111724 ago 2026
AGPL-3.024 ago 2026 · métricas 2.10.0
Packagist · npm
88Excelenteíndice de salud
eliashaeussler/cache-warmup
🔥 PHP library to warm up caches of URLs located in XML sitemaps
PHP★ 77↓ 14.1K/mes16 jul 2026
GPL-3.016 jul 2026 · métricas 2.10.0
PyPI
88Excelenteíndice de salud
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1974↓ 15.7K/mes4 sept 2026
MIT4 sept 2026 · métricas 2.10.0
crates.io
87Excelenteíndice de salud
0x676e67/wreq
An ergonomic, privacy-aware Rust HTTP Client
Rust★ 1004↓ 220.3K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
87Excelenteíndice de salud
DedSecInside/TorBot
Dark Web OSINT Tool
Python★ 472628 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3201↓ 3522/mes22 jul 2026
Licencia propia22 jul 2026 · métricas 2.10.0
crates.io
87Excelenteíndice de salud
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2669↓ 42K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm · PyPI · crates.io
87Excelenteíndice de salud
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9070/mes3 ago 2026
AGPL-3.03 ago 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
zorillajs/zorilla
Zorilla is a modular plugin framework that extends Puppeteer and Playwright with additional functionality through a clean plugin architecture. Build powerful browser automation with composable plugins for stealth mode, captcha solving, ad blocking, and much more.
TypeScript · JavaScript★ 27↓ 15.3K/mes29 jul 2026
MIT29 jul 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/mes5 ago 2026
AGPL-3.05 ago 2026 · métricas 2.10.0
Go
86Excelenteíndice de salud
openclaw/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go · Python★ 5622 ago 2026
MIT22 ago 2026 · métricas 2.10.0
NuGet
84Excelenteíndice de salud
51Degrees/device-detection-dotnet
Device detection services for 51Degrees Pipeline
C# · C++★ 1128 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
npm
83Excelenteíndice de salud
microlinkhq/top-user-agents
An always up-to-date list of the top 100 HTTP user-agents most used over the Internet.
JavaScript★ 357↓ 20.3K/mes3 ago 2026
MIT3 ago 2026 · métricas 2.10.0
Packagist
83Excelenteíndice de salud
spatie/crawler
https://spatie.be/docs/crawler
PHP★ 2829↓ 751.2K/mes20 jul 2026
MIT20 jul 2026 · métricas 2.10.0
PyPI · crates.io
81Excelenteíndice de salud
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 142313 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
z-mio/ParseHub
轻量、异步、开箱即用的社交媒体聚合解析库
Python★ 152↓ 2797/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
Packagist
80Excelenteíndice de salud
JayBizzle/Crawler-Detect
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
PHP★ 2401↓ 2.6M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0