All tags
Catalogue tag

#crawler

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

86 records
Tagged “crawler”Ranked by health index
npm
59Moderatehealth index
brandonkramer/pi-scraper
Pi extension for fast page scraping, recursive crawling, URL/site mapping, brand extraction, content diffing, PDF text extraction, and deterministic vertical extraction.
TypeScript · HTML★ 7↓ 607/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
Packagist
57Moderatehealth index
JayBizzle/Laravel-Crawler-Detect
A Laravel wrapper for CrawlerDetect - the web crawler detection library
PHP★ 324↓ 61.2K/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
PyPI
57Moderatehealth index
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/moAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
Go
56Moderatehealth index
x-way/crawlerdetect
Golang module to detect bots and crawlers via the user agent
Go★ 70Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
npm
54Moderatehealth index
codepurse/SEOCORE
Enterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health audits.
TypeScript★ 109Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
tobybgy-lsd/web-agent-runtime-bench
Local-first failure diagnosis, auto collection, repair planning, AI handoff, and verification for Playwright, crawler, RPA, and agent workflows.
Python · JavaScript★ 1↓ 3,325/moJul 26, 2026
Custom licenseJul 26, 2026 · metrics 2.10.0
npm
51Moderatehealth index
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/moJul 22, 2026
No licenseJul 22, 2026 · metrics 2.10.0
Packagist · npm
50Moderatehealth index
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
Go
50Moderatehealth index
mjc-gh/virgo
Convert a webpage into plaintext or markdown using the Chrome DevTool Protocol
Go · JavaScript★ 1Jul 17, 2026
BSD-3-ClauseJul 17, 2026 · metrics 2.10.0
Go
50Moderatehealth index
quantmind-br/repodocs
Go CLI for extracting websites, repositories, sitemaps, and package docs into structured Markdown
Go★ 0Aug 25, 2026
MITAug 25, 2026 · metrics 2.10.0
crates.io
50Moderatehealth index
spider-rs/spider_firewall
Firewall for Rust
Rust★ 1↓ 2,668/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
npm · Go · PyPI
48Weakhealth index
AlphaTechini/doc-fetch
Dynamic documentation fetching CLI that converts entire documentation sites to single markdown files for AI/LLM consumption
Svelte · Go★ 1↓ 2,547/moJul 27, 2026
No licenseJul 27, 2026 · metrics 2.10.0
Packagist
42Weakhealth index
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
Packagist
41Weakhealth index
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,078Jul 27, 2026
Custom licenseJul 27, 2026 · metrics 2.10.0
crates.io
39Weakhealth index
zekroTJA/r34-crawler
A simple CLI tool to fetch and download images from rule34.xxx
Rust★ 4Jul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
npm
36Weakhealth index
pavlealeksic/playwright-afp
Stop website fingerprinting techniques playwright edition
JavaScript★ 19↓ 4,152/moSep 2, 2026
MITSep 2, 2026 · metrics 2.10.0
NuGet
35Weakhealth index
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 62Jul 31, 2026
Custom licenseJul 31, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8KAug 12, 2026
BSD-3-ClauseAug 12, 2026 · metrics 2.10.0
npm
30At Riskhealth index
hardbulls/wbsc-crawler
No repository description published.
TypeScript★ 1↓ 2,234/moJul 25, 2026
No licenseJul 25, 2026 · metrics 2.10.0
Maven
29At Riskhealth index
code4craft/webmagic
A scalable web crawler framework for Java.
Java · HTML★ 11.7KAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm
29At Riskhealth index
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 9Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
Go
28At Riskhealth index
Synoppy/synoppy-go
Official Go SDK for Synoppy — the web-data layer for AI agents. Read, crawl, map, extract, classify & enrich any website on one key. Standard library only.
Go★ 1Jul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
Maven
23At Riskhealth index
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4KAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
22At Riskhealth index
0-EternalJunior-0/GraphCrawler
Python бібліотека для сканування веб-сайтів та побудови графу їх структури.
Python · HTML★ 1↓ 4,246/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
npm · PyPI
20At Riskhealth index
88899/gitmen-lottery
彩票 数据 双色球 大乐透 快开 超级大乐透 预测 仅供学习
Python · JavaScript★ 9Jul 27, 2026
No licenseJul 27, 2026 · metrics 2.10.0
npm
20At Riskhealth index
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0