Todas las etiquetas
Etiqueta del catálogo

#crawling

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

9 registros
Con la etiqueta «crawling»Ordenado por índice de salud
PyPI · npm
84Buenoíndice de salud
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
npm
84Buenoíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8K19 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0
PyPI
81Buenoíndice de salud
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/mes14 jul 2026
BSD-3-Clause14 jul 2026 · métricas 1.13.0
76Buenoíndice de salud
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 390017 jul 2026
MIT17 jul 2026 · métricas 1.13.0
PyPI · npm
68Moderadoíndice de salud
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9160/mes19 jul 2026
MIT19 jul 2026 · métricas 1.13.0
crates.io · npm · Packagist +1
64Moderadoíndice de salud
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 140↓ 0/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 17721 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
Packagist
46En riesgoíndice de salud
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2779/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
npm
35En riesgoíndice de salud
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 919 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0