全部标签
目录标签

#crawler

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

42 条记录
标签为“crawler”按健康指数排序
npm
84良好健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8K2026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
PyPI · npm
82良好健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0
PyPI
81良好健康指数
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/月2026年7月14日
BSD-3-Clause2026年7月14日 · 指标 1.13.0
81良好健康指数
mendableai/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 151.8K2026年7月16日
AGPL-3.02026年7月16日 · 指标 1.13.0
80良好健康指数
firecrawl/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript★ 149.9K↓ 0/月2026年7月13日
AGPL-3.02026年7月13日 · 指标 1.13.0
npm
77良好健康指数
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0
Packagist · npm
76良好健康指数
eliashaeussler/cache-warmup
🔥 PHP library to warm up caches of URLs located in XML sitemaps
PHP★ 77↓ 14.1K/月2026年7月16日
GPL-3.02026年7月16日 · 指标 1.13.0
76良好健康指数
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,9002026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm
73良好健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 1.13.0
Packagist
73良好健康指数
tomasnorre/crawler
Libraries and scripts for crawling the TYPO3 page tree. Used for re-caching, re-indexing, publishing applications etc.
PHP★ 57↓ 10.1K/月2026年7月13日
GPL-3.02026年7月13日 · 指标 1.13.0
PyPI
72良好健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,3122026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
Packagist
71良好健康指数
spatie/crawler
https://spatie.be/docs/crawler
PHP★ 2,829↓ 751.2K/月2026年7月20日
MIT2026年7月20日 · 指标 1.13.0
crates.io
71良好健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,601↓ 52.8K/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
npm
67中等健康指数
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/月2026年7月23日
MIT2026年7月23日 · 指标 1.13.0
npm
66中等健康指数
51Degrees/device-detection-node
Device detection engines for Node.js implementation of the 51Degrees Pipeline API
C++ · JavaScript★ 2↓ 254.8K/月2026年7月16日
自定义许可证2026年7月16日 · 指标 1.13.0
PyPI
65中等健康指数
z-mio/parsehub
轻量、异步、开箱即用的社交媒体聚合解析库
Python★ 125↓ 7,961/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
PyPI
64中等健康指数
PyPtt/PyPtt
The best PTT library
Python★ 726↓ 2,592/月2026年7月19日
LGPL-3.02026年7月19日 · 指标 1.13.0
PyPI
63中等健康指数
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 1772026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
Maven
62中等健康指数
fanyong920/jvppeteer
Java API For Chrome and Firefox
Java★ 8082026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
Go
61中等健康指数
openclaw/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go★ 51↓ 0/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
Go
61中等健康指数
vincentkoc/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go · Python★ 522026年7月15日
MIT2026年7月15日 · 指标 1.13.0
npm
60中等健康指数
thecodrr/fdir
⚡ The fastest directory crawler & globbing library for NodeJS. Crawls 1m files in < 1s
TypeScript★ 1,726↓ 612M/月2026年7月22日
MIT2026年7月22日 · 指标 1.13.0
npm
59中等健康指数
ODATANO/NIGHTGATE
ODATANO-NIGHTGATE connects SAP systems to the Midnight privacy chain via a CAP-based OData V4 API, enabling confidential on-chain data access and privacy-preserving transaction execution
TypeScript · JavaScript★ 3↓ 3,520/月2026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
npm
59中等健康指数
apmantza/pi-webaio
AiO webtools for Pi: Search, Fetch, Pull
TypeScript · JavaScript★ 11↓ 1,369/月2026年7月15日
MIT2026年7月15日 · 指标 1.13.0
Go
58中等健康指数
krau/manyacg
Collect, Download, Organize and Share your Favorite Anime Artworks.
Go★ 109↓ 0/月2026年7月13日
AGPL-3.02026年7月13日 · 指标 1.13.0
npm · crates.io · Go +1
58中等健康指数
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/月2026年7月16日
MIT2026年7月16日 · 指标 1.13.0
npm
57中等健康指数
mysleekdesigns/crawlforge-mcp
26-tool MCP server for web scraping, crawling, deep research & autonomous extraction — clean Markdown & structured JSON for Claude, Cursor & any MCP client. 1,000 free credits, local-Ollama LLM support.
JavaScript★ 1↓ 2,975/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
Go
57中等健康指数
uinaf/lincrawl
Local-first Linear work-graph archive CLI
Go★ 12026年7月21日
MIT2026年7月21日 · 指标 1.13.0
PyPI
55中等健康指数
bitextor/bitextor
Bitextor generates translation memories from multilingual websites
Python · Shell★ 2992026年7月21日
GPL-3.02026年7月21日 · 指标 1.13.0
Go
55中等健康指数
x-way/crawlerdetect
Golang module to detect bots and crawlers via the user agent
Go★ 702026年7月17日
MIT2026年7月17日 · 指标 1.13.0