All tags
Catalogue tag

#scraping

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

82 records
Tagged “scraping”Ranked by health index
NuGet
48Weakhealth index
darrylwhitmore/NScrape
A web scraping framework for .NET
C#★ 66Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
48Weakhealth index
sarperavci/UAForge
Generate statistically accurate User Agents and Client Hints (Sec-CH-UA). Deterministic, data-driven browser identities based on real-world market share distributions.
Python★ 9↓ 628/moAug 1, 2026
No licenseAug 1, 2026 · metrics 2.10.0
Go
47Weakhealth index
anatolykoptev/go-threads
Threads (threads.net) scraping module — public GraphQL API, LSD token auth, go-stealth TLS fingerprinting
Go★ 2Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
47Weakhealth index
online-judge-tools/api-client
API client to develop tools for competitive programming
Python★ 82Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
PyPI
47Weakhealth index
ultrafunkamsterdam/nodriver
Successor of Undetected-Chromedriver. Providing a blazing fast framework for web automation, webscraping, bots and any other creative ideas which are normally hindered by annoying anti bot systems like Captcha / CloudFlare / Imperva / hCaptcha
Python★ 4,648Aug 12, 2026
AGPL-3.0Aug 12, 2026 · metrics 2.10.0
Packagist
42Weakhealth index
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
PyPI
41Weakhealth index
DaKheera47/scraperecon
CLI recon tool for scraper developers. Detects TLS fingerprinting, JS challenges, bot protection, and rate limits across 4 stages
Python★ 42↓ 39/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
npm
39Weakhealth index
monid-ai/cli
No repository description published.
TypeScript★ 1↓ 2,563/moJul 19, 2026
No licenseJul 19, 2026 · metrics 2.10.0
npm
36Weakhealth index
pavlealeksic/playwright-afp
Stop website fingerprinting techniques playwright edition
JavaScript★ 19↓ 4,152/moSep 2, 2026
MITSep 2, 2026 · metrics 2.10.0
Go
36Weakhealth index
rusq/chromedl
Go library for scraping or downloading files bypassing Cloudflare protection and browser checks
Go★ 35Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
scrapy/parsel
Parsel lets you extract data from XML/HTML documents using XPath or CSS selectors
Python★ 1,350↓ 4.7M/moAug 13, 2026
BSD-3-ClauseAug 13, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8KAug 12, 2026
BSD-3-ClauseAug 12, 2026 · metrics 2.10.0
npm
31At Riskhealth index
ddoojoang/ttj-skills-playwright
No repository description published.
TypeScript · JavaScript★ 1↓ 3,696/moJul 29, 2026
No licenseJul 29, 2026 · metrics 2.10.0
PyPI
31At Riskhealth index
ultrafunkamsterdam/undetected-chromedriver
Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
Python★ 12.8KAug 12, 2026
GPL-3.0Aug 12, 2026 · metrics 2.10.0
Maven
29At Riskhealth index
code4craft/webmagic
A scalable web crawler framework for Java.
Java · HTML★ 11.7KAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
Go
29At Riskhealth index
syswraith/heimdall
Fast, concurrent username OSINT tool in Go — checks presence across sites via HTTP, JSON, headless browser, and OpenGraph parsing, with SQLite caching and YAML-driven site configs.
Go★ 0Jul 24, 2026
AGPL-3.0Jul 24, 2026 · metrics 2.10.0
PyPI
28At Riskhealth index
VeNoMouS/cloudscraper
A Python module to bypass Cloudflare's anti-bot page.
Python★ 6,730↓ 4.5M/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
26At Riskhealth index
scrapedatshi/scrapedatshi-py
No repository description published.
Python★ 0↓ 4,092/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
21At Riskhealth index
scrapedatshi/scrapedatshi-mcp
No repository description published.
Python★ 0↓ 2,606/moJul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
npm
20At Riskhealth index
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI
19Criticalhealth index
fake-useragent/fake-useragent
Up-to-date simple useragent faker with real world database
Python★ 4,049↓ 10.8M/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
15Criticalhealth index
binarynightowl/covid19_python
A fast, powerful, and flexible way to get up to date COVID-19 data for any major city, state, country, and total world wide data, with just one line of code
Python★ 11↓ 180/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0