全部标签
目录标签

#crawler

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

86 条记录
标签为“crawler”按健康指数排序
npm
59中等健康指数
brandonkramer/pi-scraper
Pi extension for fast page scraping, recursive crawling, URL/site mapping, brand extraction, content diffing, PDF text extraction, and deterministic vertical extraction.
TypeScript · HTML★ 7↓ 607/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
Packagist
57中等健康指数
JayBizzle/Laravel-Crawler-Detect
A Laravel wrapper for CrawlerDetect - the web crawler detection library
PHP★ 324↓ 61.2K/月2026年7月28日
MIT2026年7月28日 · 指标 2.10.0
PyPI
57中等健康指数
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/月2026年8月12日
MIT2026年8月12日 · 指标 2.10.0
Go
56中等健康指数
x-way/crawlerdetect
Golang module to detect bots and crawlers via the user agent
Go★ 702026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm
54中等健康指数
codepurse/SEOCORE
Enterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health audits.
TypeScript★ 1092026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
54中等健康指数
tobybgy-lsd/web-agent-runtime-bench
Local-first failure diagnosis, auto collection, repair planning, AI handoff, and verification for Playwright, crawler, RPA, and agent workflows.
Python · JavaScript★ 1↓ 3,325/月2026年7月26日
自定义许可证2026年7月26日 · 指标 2.10.0
npm
51中等健康指数
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/月2026年7月22日
无许可证2026年7月22日 · 指标 2.10.0
Packagist · npm
50中等健康指数
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/月2026年7月29日
MIT2026年7月29日 · 指标 2.10.0
Go
50中等健康指数
mjc-gh/virgo
Convert a webpage into plaintext or markdown using the Chrome DevTool Protocol
Go · JavaScript★ 12026年7月17日
BSD-3-Clause2026年7月17日 · 指标 2.10.0
Go
50中等健康指数
quantmind-br/repodocs
Go CLI for extracting websites, repositories, sitemaps, and package docs into structured Markdown
Go★ 02026年8月25日
MIT2026年8月25日 · 指标 2.10.0
crates.io
50中等健康指数
spider-rs/spider_firewall
Firewall for Rust
Rust★ 1↓ 2,668/月2026年7月16日
MIT2026年7月16日 · 指标 2.10.0
npm · Go · PyPI
48薄弱健康指数
AlphaTechini/doc-fetch
Dynamic documentation fetching CLI that converts entire documentation sites to single markdown files for AI/LLM consumption
Svelte · Go★ 1↓ 2,547/月2026年7月27日
无许可证2026年7月27日 · 指标 2.10.0
Packagist
42薄弱健康指数
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/月2026年7月15日
Apache-2.02026年7月15日 · 指标 2.10.0
Packagist
41薄弱健康指数
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,0782026年7月27日
自定义许可证2026年7月27日 · 指标 2.10.0
crates.io
39薄弱健康指数
zekroTJA/r34-crawler
A simple CLI tool to fetch and download images from rule34.xxx
Rust★ 42026年7月22日
MIT2026年7月22日 · 指标 2.10.0
npm
36薄弱健康指数
pavlealeksic/playwright-afp
Stop website fingerprinting techniques playwright edition
JavaScript★ 19↓ 4,152/月2026年9月2日
MIT2026年9月2日 · 指标 2.10.0
NuGet
35薄弱健康指数
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 622026年7月31日
自定义许可证2026年7月31日 · 指标 2.10.0
PyPI
34存在风险健康指数
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K2026年8月12日
BSD-3-Clause2026年8月12日 · 指标 2.10.0
npm
30存在风险健康指数
hardbulls/wbsc-crawler
该仓库未发布描述。
TypeScript★ 1↓ 2,234/月2026年7月25日
无许可证2026年7月25日 · 指标 2.10.0
Maven
29存在风险健康指数
code4craft/webmagic
A scalable web crawler framework for Java.
Java · HTML★ 11.7K2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
npm
29存在风险健康指数
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 92026年7月19日
Apache-2.02026年7月19日 · 指标 2.10.0
Go
28存在风险健康指数
Synoppy/synoppy-go
Official Go SDK for Synoppy — the web-data layer for AI agents. Read, crawl, map, extract, classify & enrich any website on one key. Standard library only.
Go★ 12026年7月28日
MIT2026年7月28日 · 指标 2.10.0
Maven
23存在风险健康指数
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4K2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
PyPI
22存在风险健康指数
0-EternalJunior-0/GraphCrawler
Python бібліотека для сканування веб-сайтів та побудови графу їх структури.
Python · HTML★ 1↓ 4,246/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
npm · PyPI
20存在风险健康指数
88899/gitmen-lottery
彩票 数据 双色球 大乐透 快开 超级大乐透 预测 仅供学习
Python · JavaScript★ 92026年7月27日
无许可证2026年7月27日 · 指标 2.10.0
npm
20存在风险健康指数
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0