All tags
Catalogue tag

#spider

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

16 records
Tagged “spider”Ranked by health index
PyPI
94Exceptionalhealth index
akfamily/akshare
AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库
Python · JavaScript★ 21.8KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
87Excellenthealth index
DedSecInside/TorBot
Dark Web OSINT Tool
Python★ 4,726Aug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
Packagist
80Excellenthealth index
JayBizzle/Crawler-Detect
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
PHP★ 2,401↓ 2.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
Maven
78Goodhealth index
Norconex/crawler
Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
Java★ 204Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
Go
67Goodhealth index
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 207Aug 2, 2026
MITAug 2, 2026 · metrics 2.10.0
npm · crates.io · Go +1
63Moderatehealth index
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
NuGet
62Moderatehealth index
zzzprojects/html-agility-pack
Html Agility Pack (HAP) is a free and open-source HTML parser written in C# to read/write DOM and supports plain XPATH or XSLT. It is a .NET code library that allows you to parse "out of the web" HTML files.
C# · HTML★ 2,846Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
moskrc/crawlerdetect
🕷CrawlerDetect is a Python library designed to identify bots, crawlers, and spiders by analyzing their user agents.
Python★ 44Jul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
Packagist
57Moderatehealth index
JayBizzle/Laravel-Crawler-Detect
A Laravel wrapper for CrawlerDetect - the web crawler detection library
PHP★ 324↓ 61.2K/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
Go
56Moderatehealth index
x-way/crawlerdetect
Golang module to detect bots and crawlers via the user agent
Go★ 70Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
48Weakhealth index
Senparc/Senparc.CO2NET
Base Common Library, support for.NET Framework &.NET Core
C#★ 365Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
Packagist
41Weakhealth index
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,078Jul 27, 2026
Custom licenseJul 27, 2026 · metrics 2.10.0
NuGet
35Weakhealth index
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 62Jul 31, 2026
Custom licenseJul 31, 2026 · metrics 2.10.0
Maven
23At Riskhealth index
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4KAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
22At Riskhealth index
0-EternalJunior-0/GraphCrawler
Python бібліотека для сканування веб-сайтів та побудови графу їх структури.
Python · HTML★ 1↓ 4,246/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0