Todas las etiquetas
Etiqueta del catálogo

#crawling

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

22 registros
Con la etiqueta «crawling»Ordenado por índice de salud
npm
98Excepcionalíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 942612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · npm
97Excepcionalíndice de salud
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K4 ago 2026
BSD-3-Clause4 ago 2026 · métricas 2.10.0
Packagist · Hex · crates.io +2
94Excepcionalíndice de salud
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3278/mes4 ago 2026
AGPL-3.04 ago 2026 · métricas 2.10.0
NuGet
92Excelenteíndice de salud
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 391628 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm · PyPI
89Excelenteíndice de salud
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 111724 ago 2026
AGPL-3.024 ago 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/mes5 ago 2026
AGPL-3.05 ago 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
linkchecker/linkchecker
check links in web documents or full websites
Python★ 1067↓ 248.6K/mes30 jul 2026
GPL-2.030 jul 2026 · métricas 2.10.0
PyPI · npm
80Excelenteíndice de salud
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9160/mes19 jul 2026
MIT19 jul 2026 · métricas 2.10.0
npm · crates.io · Packagist +1
78Buenoíndice de salud
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
77Buenoíndice de salud
tryAGI/Firecrawl
Generated C# SDK based on official Firecrawl OpenAPI specification
C#★ 42 sept 2026
MIT2 sept 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 17721 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/mes12 ago 2026
MIT12 ago 2026 · métricas 2.10.0
Packagist · npm
50Moderadoíndice de salud
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/mes29 jul 2026
MIT29 jul 2026 · métricas 2.10.0
Packagist
42Débilíndice de salud
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2779/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 2.10.0
Packagist
42Débilíndice de salud
firecrawl/firecrawl-php
El repositorio no publica descripción.
PHP★ 2↓ 2586/mes19 ago 2026
Sin licencia19 ago 2026 · métricas 2.10.0
Packagist
41Débilíndice de salud
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 207827 jul 2026
Licencia propia27 jul 2026 · métricas 2.10.0
PyPI
34En riesgoíndice de salud
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K12 ago 2026
BSD-3-Clause12 ago 2026 · métricas 2.10.0
npm
29En riesgoíndice de salud
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 919 jul 2026
Apache-2.019 jul 2026 · métricas 2.10.0
Packagist
28En riesgoíndice de salud
roach-php/core
The complete web scraping toolkit for PHP.
PHP★ 1455↓ 12K/mes27 jul 2026
Sin licencia27 jul 2026 · métricas 2.10.0
npm
20En riesgoíndice de salud
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0