全部标签
目录标签

#crawling

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

22 条记录
标签为“crawling”按健康指数排序
npm
98卓越健康指数
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI · npm
97卓越健康指数
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/月2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
NuGet
92优秀健康指数
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3,9162026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm · PyPI
89优秀健康指数
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1,1172026年8月24日
AGPL-3.02026年8月24日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
PyPI
81优秀健康指数
linkchecker/linkchecker
check links in web documents or full websites
Python★ 1,067↓ 248.6K/月2026年7月30日
GPL-2.02026年7月30日 · 指标 2.10.0
PyPI · npm
80优秀健康指数
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/月2026年7月19日
MIT2026年7月19日 · 指标 2.10.0
npm · crates.io · Packagist +1
78良好健康指数
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
77良好健康指数
tryAGI/Firecrawl
Generated C# SDK based on official Firecrawl OpenAPI specification
C#★ 42026年9月2日
MIT2026年9月2日 · 指标 2.10.0
PyPI
69良好健康指数
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 1772026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
PyPI
57中等健康指数
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/月2026年8月12日
MIT2026年8月12日 · 指标 2.10.0
Packagist · npm
50中等健康指数
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/月2026年7月29日
MIT2026年7月29日 · 指标 2.10.0
Packagist
42薄弱健康指数
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2,779/月2026年7月15日
Apache-2.02026年7月15日 · 指标 2.10.0
Packagist
42薄弱健康指数
firecrawl/firecrawl-php
该仓库未发布描述。
PHP★ 2↓ 2,586/月2026年8月19日
无许可证2026年8月19日 · 指标 2.10.0
Packagist
41薄弱健康指数
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,0782026年7月27日
自定义许可证2026年7月27日 · 指标 2.10.0
PyPI
34存在风险健康指数
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K2026年8月12日
BSD-3-Clause2026年8月12日 · 指标 2.10.0
npm
29存在风险健康指数
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 92026年7月19日
Apache-2.02026年7月19日 · 指标 2.10.0
Packagist
28存在风险健康指数
roach-php/core
The complete web scraping toolkit for PHP.
PHP★ 1,455↓ 12K/月2026年7月27日
无许可证2026年7月27日 · 指标 2.10.0
npm
20存在风险健康指数
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0