全部标签
目录标签

#web-crawling

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

6 条记录
标签为“web-crawling”按健康指数排序
PyPI · npm
84良好健康指数
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/月2026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
npm
84良好健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8K2026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
PyPI · npm
82良好健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0
PyPI
75良好健康指数
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 12.9K↓ 3M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm · crates.io · Go +1
58中等健康指数
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/月2026年7月16日
MIT2026年7月16日 · 指标 1.13.0
npm
46存在风险健康指数
NovadaLabs/novada-mcp
One MCP server for all web data — search, scrape, crawl, proxy, and AI research in a single npx install. Works with Claude, Cursor, and any MCP client.
TypeScript · HTML★ 2↓ 5,804/月2026年7月16日
无许可证2026年7月16日 · 指标 1.13.0