Todas las etiquetas
Etiqueta del catálogo

#scraper

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

82 registros
Con la etiqueta «scraper»Ordenado por índice de salud
PyPI
71Buenoíndice de salud
0xMH/pyfunda
Python API wrapper for Funda.nl, the Dutch real estate platform. Reverse-engineered mobile API client. No scraping, no Selenium, no CAPTCHA.
Python★ 17526 jul 2026
AGPL-3.026 jul 2026 · métricas 2.10.0
Go
71Buenoíndice de salud
PxyUp/fitter
New way for collect information from the API's/Websites
Go · HTML★ 13228 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
jr200-labs/polars-hist-db
(jetstream | file) --> polars dataframe <--> mariadb (bitemporal)
Python★ 1↓ 6870/mes16 jul 2026
MIT16 jul 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
sdaqo/anipy-cli
Little tool in python to watch and download anime from the terminal (the better way to watch anime). Also applicable as an API
Python★ 486↓ 23.5K/mes19 jul 2026
GPL-3.019 jul 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
stacksjs/ts-web-scraper
A powerful, type-safe web scraping library for TypeScript.
TypeScript★ 11↓ 177.5K/mes21 jul 2026
MIT21 jul 2026 · métricas 2.10.0
Maven
67Buenoíndice de salud
fanyong920/jvppeteer
Java API For Chrome and Firefox
Java★ 80817 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go
67Buenoíndice de salud
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 2072 ago 2026
MIT2 ago 2026 · métricas 2.10.0
npm · Go
67Buenoíndice de salud
pinchtab/seaportal
SeaPortal — Fast, HTTP-first content extraction for AI agents. No browser required.
Go · HTML★ 5↓ 207/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
laurengarcia/url-metadata
npm module: Fetch a url & scrape the metadata from its HTML with Node.js or the browser.
JavaScript★ 172↓ 95.4K/mes6 sept 2026
MIT6 sept 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5179/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
rocktimsaikia/meta-fetcher
Scrape metadata from a website URL
TypeScript · JavaScript★ 147↓ 3866/mes3 ago 2026
MIT3 ago 2026 · métricas 2.10.0
npm · crates.io · Go +1
63Moderadoíndice de salud
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/mes16 jul 2026
MIT16 jul 2026 · métricas 2.10.0
PyPI
62Moderadoíndice de salud
farfarfun/funflix
影视资源分享文本的结构化采集、解析与网盘链接校验 - 采集/抽取/校验分层幂等,可单独重跑
Python★ 0↓ 2003/mes30 ago 2026
MIT30 ago 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
recipe-scrapers/recipe-scrapers
A TypeScript library for scraping recipe data from cooking websites
TypeScript★ 6↓ 3693/mes2 ago 2026
MIT2 ago 2026 · métricas 2.10.0
Go
62Moderadoíndice de salud
tamnd/ccrawl-cli
A fast, friendly command line for Common Crawl: URL index search, WARC fetch, Parquet columnar queries, and dataset building.
Go★ 89 ago 2026
Apache-2.09 ago 2026 · métricas 2.10.0
Go
60Moderadoíndice de salud
AlexGustafsson/systembolaget-api
A cross-platform solution for using Systembolaget's APIs. For up-to-date data see https://github.com/AlexGustafsson/systembolaget-api-data.
Go · JavaScript★ 2826 jul 2026
Licencia propia26 jul 2026 · métricas 2.10.0
Go
60Moderadoíndice de salud
Pradumnasaraf/scrapy
Scrapy is a CLI tool used to scrape data from various websites
Go★ 429 jul 2026
Apache-2.029 jul 2026 · métricas 2.10.0
PyPI
60Moderadoíndice de salud
bpodlipnik/mirror-url
Python CLI/library for mirroring and syncing solar-physics mission data (SOHO, STEREO, PROBA-3, SolarSoft, and similar HTTP(S) data repositories), with caching, filtering, and parallel downloads
Python★ 1↓ 2582/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
Go
60Moderadoíndice de salud
man90es/BDO-REST-API
Scraper for Black Desert Online community data with a built-in API server.
Go★ 2325 jul 2026
MIT25 jul 2026 · métricas 2.10.0
Packagist
59Moderadoíndice de salud
boatracevibeproject/scraper
BVP Scraper は、ボートレースの公式サイトから出走表、直前情報、オッズ、結果をスクレイピングするための PHP ライブラリです。
PHP★ 3↓ 9665/mes18 jul 2026
MIT18 jul 2026 · métricas 2.10.0
npm
59Moderadoíndice de salud
brandonkramer/pi-scraper
Pi extension for fast page scraping, recursive crawling, URL/site mapping, brand extraction, content diffing, PDF text extraction, and deterministic vertical extraction.
TypeScript · HTML★ 7↓ 607/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
59Moderadoíndice de salud
cherchyk/MCPBrowser
El repositorio no publica descripción.
JavaScript★ 8↓ 1837/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/mes12 ago 2026
MIT12 ago 2026 · métricas 2.10.0
Go
57Moderadoíndice de salud
kwadkore/ws-scraper
[en.]ws-tcg.com scraper. Permanently forked from https://github.com/Akenaide/wsoffcli
HTML · Go★ 023 jul 2026
Apache-2.023 jul 2026 · métricas 2.10.0
PyPI
56Moderadoíndice de salud
c4road/earningspy
Your bot's best friend.
Python★ 222 ago 2026
MIT22 ago 2026 · métricas 2.10.0
Go
53Moderadoíndice de salud
Dhairya3391/kari
Kari just a tui for getting media from providers and playing it in your media player with nice to have features.
Go★ 465 sept 2026
MIT5 sept 2026 · métricas 2.10.0
Go
53Moderadoíndice de salud
tamnd/martinfowler-cli
Fetch Martin Fowler's technical articles, bliki posts, and book entries from the terminal
Go · Python★ 022 jul 2026
Apache-2.022 jul 2026 · métricas 2.10.0
Go
53Moderadoíndice de salud
tamnd/neuralnetworksdl-cli
Read Neural Networks and Deep Learning book chapters and pages as JSON from the command line
Go · Python★ 021 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
Go
53Moderadoíndice de salud
tamnd/thegradient-cli
Fetch The Gradient AI research publication posts and listings from the terminal
Go · Python★ 021 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
Go
53Moderadoíndice de salud
tamnd/theodinproject-cli
Browse The Odin Project learning paths and lessons as JSON from the command line
Go · Python★ 021 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0