全部标签
目录标签

#web-scraping

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

60 条记录
标签为“web-scraping”按健康指数排序
PyPI · npm
84良好健康指数
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 94↓ 2.5M/月2026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
npm
84良好健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 24.8K2026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
PyPI · npm
82良好健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0
PyPI
81良好健康指数
d4vinci/scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 69.4K↓ 863.2K/月2026年7月14日
BSD-3-Clause2026年7月14日 · 指标 1.13.0
81良好健康指数
mendableai/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 151.8K2026年7月16日
AGPL-3.02026年7月16日 · 指标 1.13.0
80良好健康指数
firecrawl/firecrawl
The API to search, scrape, and interact with the web at scale. 🔥
TypeScript★ 149.9K↓ 0/月2026年7月13日
AGPL-3.02026年7月13日 · 指标 1.13.0
PyPI
78良好健康指数
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6,093↓ 36.1M/月2026年7月18日
MIT2026年7月18日 · 指标 1.13.0
PyPI
75良好健康指数
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 12.9K↓ 3M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm
73良好健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 1.13.0
PyPI
72良好健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,3122026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
PyPI
72良好健康指数
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2,623/月2026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
PyPI · npm
72良好健康指数
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 152026年7月18日
MIT2026年7月18日 · 指标 1.13.0
crates.io
71良好健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,601↓ 52.8K/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
Go
70良好健康指数
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 4,7812026年7月15日
MIT2026年7月15日 · 指标 1.13.0
PyPI
70良好健康指数
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 2092026年7月20日
MIT2026年7月20日 · 指标 1.13.0
PyPI · npm
69中等健康指数
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · TypeScript · C#★ 28.5K↓ 828.1K/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm
68中等健康指数
microlinkhq/browserless
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
JavaScript★ 1,8262026年7月18日
MIT2026年7月18日 · 指标 1.13.0
npm · crates.io
68中等健康指数
olo-dot-io/uni-cli
Operations substrate for AI agents that use real software: 311 sites/tools, logged-in browsers, desktop apps, local tools, MCP, policy, evidence, AgentEnvelope v2, and self-repair.
TypeScript★ 87↓ 1,161/月2026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
npm
68中等健康指数
vmoranv/jshookmcp
js hook toolkit that all you need
TypeScript★ 1,756↓ 3,417/月2026年7月18日
AGPL-3.02026年7月18日 · 指标 1.13.0
Go · PyPI
68中等健康指数
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 422026年7月20日
MIT2026年7月20日 · 指标 1.13.0
Go
67中等健康指数
HappyHackingSpace/dit
HTML page, form and field type classifier using ML (LogReg + CRF)
Go★ 172026年7月22日
MIT2026年7月22日 · 指标 1.13.0
npm
67中等健康指数
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/月2026年7月23日
MIT2026年7月23日 · 指标 1.13.0
crates.io
66中等健康指数
0x676e67/wreq-util
Common utilities for wreq
Rust★ 88↓ 221.4K/月2026年7月14日
Apache-2.02026年7月14日 · 指标 1.13.0
PyPI
65中等健康指数
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 154↓ 13.4M/月2026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
npm
64中等健康指数
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2,900/月2026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
crates.io
64中等健康指数
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 9↓ 0/月2026年7月14日
Apache-2.02026年7月14日 · 指标 1.13.0
crates.io · npm · Packagist +1
64中等健康指数
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 140↓ 0/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
Go · npm
62中等健康指数
1broseidon/ketch
Fast, stateless CLI for web search and scrape. Built for AI agents.
Go★ 3882026年7月18日
MIT2026年7月18日 · 指标 1.13.0
Go
62中等健康指数
ChristopherDavenport/unblink
Pure-Go browser for AI: fetch pages, optionally render JS (no Chromium), get clean token-budgeted Markdown over MCP
Go★ 72026年7月22日
MIT2026年7月22日 · 指标 1.13.0
Go · npm
62中等健康指数
felixgeelhaar/scout
Browser automation, one binary. The simpler alternative to Playwright — no Node, no Python, no runtime. Library, CLI, MCP server, and chat UI for any AI agent.
Go★ 62026年7月15日
MIT2026年7月15日 · 指标 1.13.0