Todas las etiquetas
Etiqueta del catálogo

#scraping

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

82 registros
Con la etiqueta «scraping»Ordenado por índice de salud
NuGet
48Débilíndice de salud
darrylwhitmore/NScrape
A web scraping framework for .NET
C#★ 6617 jul 2026
MIT17 jul 2026 · métricas 2.10.0
PyPI
48Débilíndice de salud
sarperavci/UAForge
Generate statistically accurate User Agents and Client Hints (Sec-CH-UA). Deterministic, data-driven browser identities based on real-world market share distributions.
Python★ 9↓ 628/mes1 ago 2026
Sin licencia1 ago 2026 · métricas 2.10.0
Go
47Débilíndice de salud
anatolykoptev/go-threads
Threads (threads.net) scraping module — public GraphQL API, LSD token auth, go-stealth TLS fingerprinting
Go★ 213 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
47Débilíndice de salud
online-judge-tools/api-client
API client to develop tools for competitive programming
Python★ 8215 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI
47Débilíndice de salud
ultrafunkamsterdam/nodriver
Successor of Undetected-Chromedriver. Providing a blazing fast framework for web automation, webscraping, bots and any other creative ideas which are normally hindered by annoying anti bot systems like Captcha / CloudFlare / Imperva / hCaptcha
Python★ 464812 ago 2026
AGPL-3.012 ago 2026 · métricas 2.10.0
Packagist
42Débilíndice de salud
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2779/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 2.10.0
PyPI
41Débilíndice de salud
DaKheera47/scraperecon
CLI recon tool for scraper developers. Detects TLS fingerprinting, JS challenges, bot protection, and rate limits across 4 stages
Python★ 42↓ 39/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
npm
39Débilíndice de salud
monid-ai/cli
El repositorio no publica descripción.
TypeScript★ 1↓ 2563/mes19 jul 2026
Sin licencia19 jul 2026 · métricas 2.10.0
npm
36Débilíndice de salud
pavlealeksic/playwright-afp
Stop website fingerprinting techniques playwright edition
JavaScript★ 19↓ 4152/mes2 sept 2026
MIT2 sept 2026 · métricas 2.10.0
Go
36Débilíndice de salud
rusq/chromedl
Go library for scraping or downloading files bypassing Cloudflare protection and browser checks
Go★ 355 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
34En riesgoíndice de salud
scrapy/parsel
Parsel lets you extract data from XML/HTML documents using XPath or CSS selectors
Python★ 1350↓ 4.7M/mes13 ago 2026
BSD-3-Clause13 ago 2026 · métricas 2.10.0
PyPI
34En riesgoíndice de salud
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K12 ago 2026
BSD-3-Clause12 ago 2026 · métricas 2.10.0
npm
31En riesgoíndice de salud
ddoojoang/ttj-skills-playwright
El repositorio no publica descripción.
TypeScript · JavaScript★ 1↓ 3696/mes29 jul 2026
Sin licencia29 jul 2026 · métricas 2.10.0
PyPI
31En riesgoíndice de salud
ultrafunkamsterdam/undetected-chromedriver
Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
Python★ 12.8K12 ago 2026
GPL-3.012 ago 2026 · métricas 2.10.0
Maven
29En riesgoíndice de salud
code4craft/webmagic
A scalable web crawler framework for Java.
Java · HTML★ 11.7K27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
Go
29En riesgoíndice de salud
syswraith/heimdall
Fast, concurrent username OSINT tool in Go — checks presence across sites via HTTP, JSON, headless browser, and OpenGraph parsing, with SQLite caching and YAML-driven site configs.
Go★ 024 jul 2026
AGPL-3.024 jul 2026 · métricas 2.10.0
PyPI
28En riesgoíndice de salud
VeNoMouS/cloudscraper
A Python module to bypass Cloudflare's anti-bot page.
Python★ 6730↓ 4.5M/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
26En riesgoíndice de salud
scrapedatshi/scrapedatshi-py
El repositorio no publica descripción.
Python★ 0↓ 4092/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
PyPI
21En riesgoíndice de salud
scrapedatshi/scrapedatshi-mcp
El repositorio no publica descripción.
Python★ 0↓ 2606/mes18 jul 2026
MIT18 jul 2026 · métricas 2.10.0
npm
20En riesgoíndice de salud
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
19Críticoíndice de salud
fake-useragent/fake-useragent
Up-to-date simple useragent faker with real world database
Python★ 4049↓ 10.8M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
15Críticoíndice de salud
binarynightowl/covid19_python
A fast, powerful, and flexible way to get up to date COVID-19 data for any major city, state, country, and total world wide data, with just one line of code
Python★ 11↓ 180/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0