全部标签
目录标签

#scraper

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

82 条记录
标签为“scraper”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
npm
94卓越健康指数
cheeriojs/cheerio
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
TypeScript · HTML · Astro★ 30.4K↓ 109M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm
91优秀健康指数
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
PyPI · npm
91优秀健康指数
henrique-coder/perplexity-webui-scraper
An advanced, high-performance Python client, MCP server, and REST API for reverse-engineering Perplexity AI's WebUI.
Python★ 115↓ 21.8K/月2026年8月18日
MIT2026年8月18日 · 指标 2.10.0
PyPI · npm
90优秀健康指数
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2,566↓ 2.2M/月2026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
PyPI
90优秀健康指数
mikf/gallery-dl
Command-line program to download image galleries and collections from several image hosting sites
Python★ 19.1K2026年8月5日
GPL-2.02026年8月5日 · 指标 2.10.0
Maven
89优秀健康指数
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
Java★ 1,9362026年8月9日
GPL-3.02026年8月9日 · 指标 2.10.0
crates.io
87优秀健康指数
0x676e67/wreq
An ergonomic, privacy-aware Rust HTTP Client
Rust★ 1,004↓ 220.3K/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
crates.io
87优秀健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm · PyPI · crates.io
87优秀健康指数
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/月2026年8月3日
AGPL-3.02026年8月3日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
RubyGems
83优秀健康指数
html2rss/html2rss
📰 Build RSS 2.0 feeds from websites (and JSON APIs) automatically or with a few CSS selectors.
Ruby★ 1632026年8月11日
MIT2026年8月11日 · 指标 2.10.0
npm · RubyGems
83优秀健康指数
huginn/huginn
Create agents that monitor and act on your behalf. Your agents are standing by!
Ruby★ 49.7K2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI
83优秀健康指数
jordantete/OddsHarvester
A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯
Python · HTML★ 2092026年7月20日
MIT2026年7月20日 · 指标 2.10.0
PyPI · crates.io
81优秀健康指数
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 1,4232026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
npm
81优秀健康指数
Xquik-dev/tweetclaw
OpenClaw plugin to search tweets, search replies, post tweets, export followers, manage media, monitor X/Twitter, and run giveaway draws via Xquik. Not affiliated with X Corp.
TypeScript · JavaScript★ 89↓ 1,243/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm
78良好健康指数
MrAdex77/google-play-scraper
Google Play scraper for Node.js with a fully typed TypeScript API. App details, search, top charts, reviews, permissions and data safety. A modern replacement for the unmaintained google-play-scraper.
TypeScript★ 5↓ 4,236/月2026年7月23日
MIT2026年7月23日 · 指标 2.10.0
PyPI
78良好健康指数
rushter/selectolax
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Cython · Python★ 1,6562026年7月20日
MIT2026年7月20日 · 指标 2.10.0
npm
78良好健康指数
zerodytrash/TikTok-Live-Connector
Node.js library to receive live stream events (comments, gifts, etc.) in realtime from TikTok LIVE.
TypeScript★ 2,080↓ 110.9K/月2026年7月21日
AGPL-3.02026年7月21日 · 指标 2.10.0
Go
75良好健康指数
Anastylosis/FSS
fss or Full Studio Scraper - scrape all videos metadata of your favourite performers or studios
Go★ 72026年8月22日
GPL-3.02026年8月22日 · 指标 2.10.0
npm
75良好健康指数
TVScoundrel/agentforge
该仓库未发布描述。
TypeScript★ 1↓ 21.9K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
PyPI
75良好健康指数
bdperkin/gamesheet-sdk-py
Unofficial Python SDK and CLI for the GameSheet Inc. platform — automates WebUI procedures where a native API is absent.
Python★ 0↓ 2,468/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
NuGet
75良好健康指数
soenneker/soenneker.playwrights.crawler
A configurable Playwright crawler with rich stealth and control options.
C#★ 02026年8月31日
MIT2026年8月31日 · 指标 2.10.0
Go · npm
73良好健康指数
diadata-org/decentral-feeder
Feeder node for the Lumina oracle network, providing trustless and verifiable on-chain data.
Solidity★ 42026年7月22日
GPL-3.02026年7月22日 · 指标 2.10.0
npm
73良好健康指数
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/月2026年7月26日
无许可证2026年7月26日 · 指标 2.10.0