全部标签
目录标签

#web-scraping

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

105 条记录
标签为“web-scraping”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI · npm
97卓越健康指数
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/月2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
Packagist · PyPI
96卓越健康指数
php-curl-class/php-curl-class
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
PHP★ 3,301↓ 140.4K/月2026年7月20日
Unlicense2026年7月20日 · 指标 2.10.0
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
lexiforest/curl_cffi
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Python★ 6,395↓ 45M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
npm
93卓越健康指数
browserbase/stagehand
The SDK For Browser Agents
TypeScript · MDX★ 23.7K↓ 4.8M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
npm · Go
92优秀健康指数
pinchtab/pinchtab
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Go · Shell★ 10.1K↓ 7,120/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
91优秀健康指数
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
Maven
90优秀健康指数
jhy/jsoup
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Java · HTML★ 11.4K2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
npm
89优秀健康指数
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7,332↓ 492.7K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm · PyPI
88优秀健康指数
MODSetter/SurfSense
Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9
Python · TypeScript★ 15.8K2026年8月5日
自定义许可证2026年8月5日 · 指标 2.10.0
PyPI
88优秀健康指数
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/月2026年9月4日
MIT2026年9月4日 · 指标 2.10.0
npm
87优秀健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 2.10.0
npm
87优秀健康指数
jo-inc/camofox-browser
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8,924↓ 243.1K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
crates.io
87优秀健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm · PyPI · crates.io
87优秀健康指数
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/月2026年8月3日
AGPL-3.02026年8月3日 · 指标 2.10.0
PyPI
86优秀健康指数
flytohub/flyto-core
Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 451 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.
Python★ 473↓ 2,623/月2026年7月19日
Apache-2.02026年7月19日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
Go
86优秀健康指数
gosom/google-maps-scraper
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place
Go · HTML★ 5,6412026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm
86优秀健康指数
kepano/defuddle
Get the main content of any page as Markdown.
TypeScript★ 9,191↓ 2.2M/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
PyPI · npm
86优秀健康指数
n24q02m/wet-mcp
Open-source MCP server for AI agents: web search, content extraction, and library docs -- 5-strategy scraping, runs without API keys.
Python★ 152026年7月18日
MIT2026年7月18日 · 指标 2.10.0
PyPI
83优秀健康指数
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 2092026年7月20日
MIT2026年7月20日 · 指标 2.10.0
PyPI · npm
83优秀健康指数
sportsdataverse/sportsdataverse-py
sportsdataverse python package
Python★ 111↓ 15.8K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
PyPI · crates.io
81优秀健康指数
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 1,4232026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
PyPI
81优秀健康指数
ArchiveBox/abx-plugins
🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/月2026年7月27日
MIT2026年7月27日 · 指标 2.10.0