全部标签
目录标签

#web-crawler

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

21 条记录
标签为“web-crawler”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
npm
89优秀健康指数
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7,332↓ 492.7K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm · PyPI
89优秀健康指数
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,1172026年8月24日
AGPL-3.02026年8月24日 · 指标 2.10.0
npm
87优秀健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 2.10.0
crates.io
87优秀健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm · PyPI · crates.io
87优秀健康指数
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/月2026年8月3日
AGPL-3.02026年8月3日 · 指标 2.10.0
npm
80优秀健康指数
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 1,8262026年7月18日
MIT2026年7月18日 · 指标 2.10.0
Maven
78良好健康指数
Norconex/crawler
Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
Java★ 2042026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
Go
78良好健康指数
kinorai/omnifeed
LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible
Go★ 62026年9月3日
MIT2026年9月3日 · 指标 2.10.0
npm · crates.io · Packagist +1
78良好健康指数
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
npm · crates.io
75良好健康指数
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/月2026年7月26日
AGPL-3.02026年7月26日 · 指标 2.10.0
npm · PyPI · crates.io
75良好健康指数
Goldziher/basemind
Full AI context and content layer for coding agents over one MCP server — tree-sitter code-map, document RAG, shared memory, multi-agent comms, web crawl, git history + blame. 300+ languages, 10+ agent harnesses, pure Rust.
Rust★ 69↓ 13.6K/月2026年7月27日
MIT2026年7月27日 · 指标 2.10.0
Go
67良好健康指数
gosom/scrapemate
Golang Crawling and scraping framework
Go★ 2072026年8月2日
MIT2026年8月2日 · 指标 2.10.0
npm
63中等健康指数
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
npm · crates.io · Go +1
63中等健康指数
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/月2026年7月16日
MIT2026年7月16日 · 指标 2.10.0
npm
41薄弱健康指数
Aegis-Runner/AegisRunner
AegisRunner crawls any website from a single URL, discovers every page, form, and interaction, then uses AI to generate a complete Playwright test suite. Also runs accessibility, SEO, security, and performance audits on every page — no recording, no scripting, no setup.
JavaScript★ 0↓ 8,920/月2026年7月25日
MIT2026年7月25日 · 指标 2.10.0
Maven
23存在风险健康指数
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4K2026年8月27日
MIT2026年8月27日 · 指标 2.10.0