全部标签
目录标签

#crawler

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

85 条记录
标签为“crawler”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
NuGet
92优秀健康指数
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,9162026年8月28日
MIT2026年8月28日 · 指标 2.10.0
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm
91优秀健康指数
apify/apify-sdk-js
Apify SDK monorepo
MDX · TypeScript★ 180↓ 149.3K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
PyPI · npm
90优秀健康指数
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
TypeScript · JavaScript★ 2,566↓ 2.2M/月2026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
PyPI
90优秀健康指数
mikf/gallery-dl
Command-line program to download image galleries and collections from several image hosting sites
Python★ 19.1K2026年8月5日
GPL-2.02026年8月5日 · 指标 2.10.0
Maven
89优秀健康指数
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
Java★ 1,9362026年8月9日
GPL-3.02026年8月9日 · 指标 2.10.0
Packagist
89优秀健康指数
tomasnorre/crawler
Libraries and scripts for crawling the TYPO3 page tree. Used for re-caching, re-indexing, publishing applications etc.
PHP★ 57↓ 8,254/月2026年8月22日
GPL-3.02026年8月22日 · 指标 2.10.0
npm · PyPI
89优秀健康指数
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,1172026年8月24日
AGPL-3.02026年8月24日 · 指标 2.10.0
Packagist · npm
88优秀健康指数
eliashaeussler/cache-warmup
🔥 PHP library to warm up caches of URLs located in XML sitemaps
PHP★ 77↓ 14.1K/月2026年7月16日
GPL-3.02026年7月16日 · 指标 2.10.0
PyPI
88优秀健康指数
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/月2026年9月4日
MIT2026年9月4日 · 指标 2.10.0
crates.io
87优秀健康指数
0x676e67/wreq
An ergonomic, privacy-aware Rust HTTP Client
Rust★ 1,004↓ 220.3K/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
87优秀健康指数
DedSecInside/TorBot
Dark Web OSINT Tool
Python★ 4,7262026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
npm
87优秀健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 2.10.0
crates.io
87优秀健康指数
spider-rs/spider
Get web data for AI agents and LLMs
Rust★ 2,669↓ 42K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm · PyPI · crates.io
87优秀健康指数
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/月2026年8月3日
AGPL-3.02026年8月3日 · 指标 2.10.0
npm
87优秀健康指数
zorillajs/zorilla
Zorilla is a modular plugin framework that extends Puppeteer and Playwright with additional functionality through a clean plugin architecture. Build powerful browser automation with composable plugins for stealth mode, captcha solving, ad blocking, and much more.
TypeScript · JavaScript★ 27↓ 15.3K/月2026年7月29日
MIT2026年7月29日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
Go
86优秀健康指数
openclaw/crawlkit
Shared Go infrastructure for local-first crawler archives.
Go · Python★ 562026年8月22日
MIT2026年8月22日 · 指标 2.10.0
NuGet
84优秀健康指数
51Degrees/device-detection-dotnet
Device detection services for 51Degrees Pipeline
C# · C++★ 112026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
npm
83优秀健康指数
microlinkhq/top-user-agents
An always up-to-date list of the top 100 HTTP user-agents most used over the Internet.
JavaScript★ 357↓ 20.3K/月2026年8月3日
MIT2026年8月3日 · 指标 2.10.0
Packagist
83优秀健康指数
spatie/crawler
https://spatie.be/docs/crawler
PHP★ 2,829↓ 751.2K/月2026年7月20日
MIT2026年7月20日 · 指标 2.10.0
PyPI · crates.io
81优秀健康指数
0x676e67/wreq-python
An ergonomic, privacy-aware Python HTTP Client
Rust · Python★ 1,4232026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
PyPI
81优秀健康指数
z-mio/ParseHub
轻量、异步、开箱即用的社交媒体聚合解析库
Python★ 152↓ 2,797/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
Packagist
80优秀健康指数
JayBizzle/Crawler-Detect
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
PHP★ 2,401↓ 2.6M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0