全部标签
目录标签

#web-crawling

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

9 条记录
标签为“web-crawling”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI · npm
97卓越健康指数
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/月2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
PyPI
91优秀健康指数
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
npm
80优秀健康指数
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm
67良好健康指数
superagents-lab/search1api-mcp
Official Search1API MCP server for web search, news, crawling, sitemaps, and trends—hosted with OAuth 2.1 or local via npm.
TypeScript · JavaScript★ 173↓ 2,722/月2026年8月24日
MIT2026年8月24日 · 指标 2.10.0
npm · crates.io · Go +1
63中等健康指数
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/月2026年7月16日
MIT2026年7月16日 · 指标 2.10.0
npm
47薄弱健康指数
NovadaLabs/novada-mcp
One MCP server for all web data — search, scrape, crawl, proxy, and AI research in a single npx install. Works with Claude, Cursor, and any MCP client.
TypeScript · HTML★ 2↓ 5,804/月2026年7月16日
无许可证2026年7月16日 · 指标 2.10.0