全部标签
目录标签

#webscraping

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

19 条记录
标签为“webscraping”按健康指数排序
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
assafelovic/gpt-researcher
An autonomous agent that conducts deep research on any data using any LLM providers
Python · TypeScript★ 28.8K2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
91优秀健康指数
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
PyPI · npm
87优秀健康指数
daijro/camoufox
🦊 Anti-detect browser
C++ · Python · JavaScript★ 11.7K↓ 851.4K/月2026年9月6日
MPL-2.02026年9月6日 · 指标 2.10.0
npm
87优秀健康指数
openzim/node-libzim
Libzim binding for Node.js: read/write ZIM files in Javascript
C++ · TypeScript★ 35↓ 3,078/月2026年7月26日
GPL-3.02026年7月26日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
npm · RubyGems
83优秀健康指数
huginn/huginn
Create agents that monitor and act on your behalf. Your agents are standing by!
Ruby★ 49.7K2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI
81优秀健康指数
cdpdriver/zendriver
A blazing fast, async-first, undetectable webscraping/web automation framework based on ultrafunkamsterdam/nodriver. Now with Docker support!
Python★ 1,3752026年7月27日
AGPL-3.02026年7月27日 · 指标 2.10.0
PyPI · npm
80优秀健康指数
CloakHQ/CloakBrowser
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI
78良好健康指数
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 155↓ 15M/月2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
PyPI
63中等健康指数
clearcotelabs/clearcote-browser
Open-source stealth Chromium 149 with engine-level fingerprint spoofing - de-Googled, drop-in Playwright, fully buildable and verifiable from source.
Python · TypeScript · C#★ 31↓ 141/月2026年7月18日
BSD-3-Clause2026年7月18日 · 指标 2.10.0
PyPI
57中等健康指数
Kaliiiiiiiiii-Vinyzu/patchright-python
Undetected Python version of the Playwright testing and automation library.
Python★ 1,4332026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
npm
51中等健康指数
Hyper-Solutions/hyper-sdk-js
JavaScript / TypeScript SDK for Bot Protection Bypass - Automate Akamai, Incapsula, Kasada, and DataDome. No browsers required. Solve challenges and generate valid sensors/cookies via API.
TypeScript★ 54↓ 31.6K/月2026年7月21日
MIT2026年7月21日 · 指标 2.10.0
npm
51中等健康指数
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/月2026年7月22日
无许可证2026年7月22日 · 指标 2.10.0
Go
50中等健康指数
Hyper-Solutions/hyper-sdk-go
Go SDK for Bot Protection Bypass - Automate Akamai, Incapsula, Kasada, and DataDome. No browsers required. Solve challenges and generate valid sensors/cookies via API.
Go★ 622026年7月26日
MIT2026年7月26日 · 指标 2.10.0
PyPI
34存在风险健康指数
requests-cache/requests-cache
Persistent HTTP cache for python requests
Python★ 1,4992026年7月20日
BSD-2-Clause2026年7月20日 · 指标 2.10.0