All tags
Catalogue tag

#crawler

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

86 records
Tagged “crawler”Ranked by health index
npm
80Excellenthealth index
NeuraLegion/bright-cli
Command Line Interface (CLI) tool for BrightSec's solutions.
TypeScript★ 22↓ 3,745/moJul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
npm
78Goodhealth index
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
Maven
78Goodhealth index
Norconex/crawler
Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
Java★ 204Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
rushter/selectolax
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 1,656Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
npm
77Goodhealth index
51Degrees/device-detection-node
Device detection engines for Node.js implementation of the 51Degrees Pipeline API
C++ · JavaScript★ 2↓ 254.8K/moJul 16, 2026
Custom licenseJul 16, 2026 · metrics 2.10.0
NuGet
77Goodhealth index
JaCraig/Spidey
A multi threaded web crawler library that is generic enough to allow different engines to be swapped in.
C#★ 16Aug 21, 2026
Apache-2.0Aug 21, 2026 · metrics 2.10.0
Packagist
77Goodhealth index
spatie/robots-txt
Determine if a page may be crawled from robots.txt, robots meta tags and robot headers
PHP★ 258↓ 785.7K/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
npm · crates.io
75Goodhealth index
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/moJul 26, 2026
AGPL-3.0Jul 26, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
Liu233w/ojhunt-lite
A lightweight async Python tool for querying Online Judge (OJ) statistics across multiple platforms. Track your accepted problems (AC) and total submissions from 29+ competitive programming platforms.
Python★ 5↓ 3,655/moJul 29, 2026
BSD-2-ClauseJul 29, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
PyPtt/PyPtt
The best PTT library
Python★ 726↓ 2,592/moJul 19, 2026
LGPL-3.0Jul 19, 2026 · metrics 2.10.0
npm
73Goodhealth index
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/moJul 26, 2026
No licenseJul 26, 2026 · metrics 2.10.0
npm
73Goodhealth index
iannuttall/seo
The only SEO skill your agent needs. 70+ SEO audit tools through a local CLI and MCP server, using your own crawl, Search Console, and GA4 data.
TypeScript · Astro★ 81↓ 4,913/moAug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
npm
71Goodhealth index
qualweb/qualweb
No repository description published.
TypeScript · HTML · CSS★ 23↓ 14.6K/moSep 5, 2026
ISCSep 5, 2026 · metrics 2.10.0
Go
71Goodhealth index
vincentkoc/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go · Python★ 52Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 177Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
Maven
67Goodhealth index
fanyong920/jvppeteer
Java API For Chrome and Firefox
Java★ 808Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go
67Goodhealth index
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 207Aug 2, 2026
MITAug 2, 2026 · metrics 2.10.0
Go
67Goodhealth index
krau/ManyACG
Collect, Download, Organize and Share your Favorite Anime Artworks.
Go★ 115Aug 22, 2026
AGPL-3.0Aug 22, 2026 · metrics 2.10.0
npm
65Goodhealth index
cosmocoder/mcp-web-docs
Self-hosted MCP server to index and search any documentation site—public or private
TypeScript★ 1↓ 4,014/moJul 26, 2026
MITJul 26, 2026 · metrics 2.10.0
npm
63Moderatehealth index
ODATANO/NIGHTGATE
ODATANO-NIGHTGATE connects SAP systems to the Midnight privacy chain via a CAP-based OData V4 API, enabling confidential on-chain data access and privacy-preserving transaction execution
TypeScript · JavaScript★ 3↓ 3,520/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
npm
63Moderatehealth index
apmantza/pi-webaio
AiO webtools for Pi: Search, Fetch, Pull
TypeScript · JavaScript★ 11↓ 1,369/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
RubyGems
63Moderatehealth index
loadkpi/crawler_detect
Ruby gem to detect bots and crawlers via the user agent
Ruby★ 148Sep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
npm · crates.io · Go +1
63Moderatehealth index
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
npm
63Moderatehealth index
thecodrr/fdir
⚡ The fastest directory crawler & globbing library for NodeJS. Crawls 1m files in < 1s
TypeScript★ 1,728↓ 630M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
Go
63Moderatehealth index
uinaf/lincrawl
Local-first Linear work-graph archive CLI
Go★ 1Jul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
NuGet
62Moderatehealth index
zzzprojects/html-agility-pack
Html Agility Pack (HAP) is a free and open-source HTML parser written in C# to read/write DOM and supports plain XPATH or XSLT. It is a .NET code library that allows you to parse "out of the web" HTML files.
C# · HTML★ 2,846Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
moskrc/crawlerdetect
🕷CrawlerDetect is a Python library designed to identify bots, crawlers, and spiders by analyzing their user agents.
Python★ 44Jul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
bitextor/bitextor
Bitextor generates translation memories from multilingual websites
Python · Shell★ 299Jul 21, 2026
GPL-3.0Jul 21, 2026 · metrics 2.10.0