Todas las etiquetas
Etiqueta del catálogo

#scraper

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

82 registros
Con la etiqueta «scraper»Ordenado por índice de salud
npm
98Excepcionalíndice de salud
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 942612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K4 ago 2026
BSD-3-Clause4 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 17320 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
npm
94Excepcionalíndice de salud
cheeriojs/cheerio
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
TypeScript · HTML · Astro★ 30.4K↓ 109M/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
Packagist · Hex · crates.io +2
94Excepcionalíndice de salud
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3278/mes4 ago 2026
AGPL-3.04 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6718↓ 14M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm
91Excelenteíndice de salud
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
PyPI · npm
91Excelenteíndice de salud
henrique-coder/perplexity-webui-scraper
An advanced, high-performance Python client, MCP server, and REST API for reverse-engineering Perplexity AI's WebUI.
Python★ 115↓ 21.8K/mes18 ago 2026
MIT18 ago 2026 · métricas 2.10.0
PyPI · npm
90Excelenteíndice de salud
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2566↓ 2.2M/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
mikf/gallery-dl
Command-line program to download image galleries and collections from several image hosting sites
Python★ 19.1K5 ago 2026
GPL-2.05 ago 2026 · métricas 2.10.0
Maven
89Excelenteíndice de salud
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
Java★ 19369 ago 2026
GPL-3.09 ago 2026 · métricas 2.10.0
crates.io
87Excelenteíndice de salud
0x676e67/wreq
An ergonomic, privacy-aware Rust HTTP Client
Rust★ 1004↓ 220.3K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
crates.io
87Excelenteíndice de salud
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2669↓ 42K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm · PyPI · crates.io
87Excelenteíndice de salud
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9070/mes3 ago 2026
AGPL-3.03 ago 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/mes5 ago 2026
AGPL-3.05 ago 2026 · métricas 2.10.0
RubyGems
83Excelenteíndice de salud
html2rss/html2rss
📰 Build RSS 2.0 feeds from websites (and JSON APIs) automatically or with a few CSS selectors.
Ruby★ 16311 ago 2026
MIT11 ago 2026 · métricas 2.10.0
npm · RubyGems
83Excelenteíndice de salud
huginn/huginn
Create agents that monitor and act on your behalf. Your agents are standing by!
Ruby★ 49.7K4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 20920 jul 2026
MIT20 jul 2026 · métricas 2.10.0
PyPI · crates.io
81Excelenteíndice de salud
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 142313 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
npm
81Excelenteíndice de salud
Xquik-dev/tweetclaw
OpenClaw plugin to search tweets, search replies, post tweets, export followers, manage media, monitor X/Twitter, and run giveaway draws via Xquik. Not affiliated with X Corp.
TypeScript · JavaScript★ 89↓ 1243/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4236/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
rushter/selectolax
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 165620 jul 2026
MIT20 jul 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
zerodytrash/TikTok-Live-Connector
Node.js library to receive live stream events (comments, gifts, etc.) in realtime from TikTok LIVE.
TypeScript★ 2080↓ 110.9K/mes21 jul 2026
AGPL-3.021 jul 2026 · métricas 2.10.0
Go
75Buenoíndice de salud
Anastylosis/FSS
fss or Full Studio Scraper - scrape all videos metadata of your favourite performers or studios
Go★ 722 ago 2026
GPL-3.022 ago 2026 · métricas 2.10.0
npm
75Buenoíndice de salud
TVScoundrel/agentforge
El repositorio no publica descripción.
TypeScript★ 1↓ 21.9K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
PyPI
75Buenoíndice de salud
bdperkin/gamesheet-sdk-py
Unofficial Python SDK and CLI for the GameSheet Inc. platform — automates WebUI procedures where a native API is absent.
Python★ 0↓ 2468/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
Go · npm
73Buenoíndice de salud
diadata-org/decentral-feeder
Feeder node for the Lumina oracle network, providing trustless and verifiable on-chain data.
Solidity★ 422 jul 2026
GPL-3.022 jul 2026 · métricas 2.10.0
npm
73Buenoíndice de salud
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/mes26 jul 2026
Sin licencia26 jul 2026 · métricas 2.10.0