Alle Tags
Katalog-Tag

#crawling

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

22 Einträge
Getaggt als „crawling“Geordnet nach Gesundheitsindex
npm
98AußergewöhnlichGesundheitsindex
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
PyPI · npm
98AußergewöhnlichGesundheitsindex
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9.42612. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI · npm
97AußergewöhnlichGesundheitsindex
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/Monat27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K4. Aug. 2026
BSD-3-Clause4. Aug. 2026 · Metriken 2.10.0
Packagist · Hex · crates.io +2
94AußergewöhnlichGesundheitsindex
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3.278/Monat4. Aug. 2026
AGPL-3.04. Aug. 2026 · Metriken 2.10.0
NuGet
92ExzellentGesundheitsindex
hardkoded/puppeteer-sharp
Headless Chrome .NET API
C#★ 3.91628. Aug. 2026
MIT28. Aug. 2026 · Metriken 2.10.0
npm · PyPI
89ExzellentGesundheitsindex
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
TypeScript★ 1.11724. Aug. 2026
AGPL-3.024. Aug. 2026 · Metriken 2.10.0
npm
86ExzellentGesundheitsindex
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/Monat5. Aug. 2026
AGPL-3.05. Aug. 2026 · Metriken 2.10.0
PyPI
81ExzellentGesundheitsindex
linkchecker/linkchecker
check links in web documents or full websites
Python★ 1.067↓ 248.6K/Monat30. Juli 2026
GPL-2.030. Juli 2026 · Metriken 2.10.0
PyPI · npm
80ExzellentGesundheitsindex
ArchiveBox/abx-dl
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9.160/Monat19. Juli 2026
MIT19. Juli 2026 · Metriken 2.10.0
npm · crates.io · Packagist +1
78GutGesundheitsindex
xberg-io/crawlberg
High-performance web crawling engine with bindings for 11 languages
Rust★ 149↓ 685/Monat1. Aug. 2026
MIT1. Aug. 2026 · Metriken 2.10.0
77GutGesundheitsindex
tryAGI/Firecrawl
Generated C# SDK based on official Firecrawl OpenAPI specification
C#★ 42. Sept. 2026
MIT2. Sept. 2026 · Metriken 2.10.0
PyPI
69GutGesundheitsindex
adbar/courlan
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Python★ 17721. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
PyPI
57MittelGesundheitsindex
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/Monat12. Aug. 2026
MIT12. Aug. 2026 · Metriken 2.10.0
Packagist · npm
50MittelGesundheitsindex
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/Monat29. Juli 2026
MIT29. Juli 2026 · Metriken 2.10.0
Packagist
42SchwachGesundheitsindex
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2.779/Monat15. Juli 2026
Apache-2.015. Juli 2026 · Metriken 2.10.0
Packagist
42SchwachGesundheitsindex
firecrawl/firecrawl-php
Keine Repository-Beschreibung veröffentlicht.
PHP★ 2↓ 2.586/Monat19. Aug. 2026
Keine Lizenz19. Aug. 2026 · Metriken 2.10.0
Packagist
41SchwachGesundheitsindex
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2.07827. Juli 2026
Eigene Lizenz27. Juli 2026 · Metriken 2.10.0
PyPI
34GefährdetGesundheitsindex
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K12. Aug. 2026
BSD-3-Clause12. Aug. 2026 · Metriken 2.10.0
npm
29GefährdetGesundheitsindex
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 919. Juli 2026
Apache-2.019. Juli 2026 · Metriken 2.10.0
Packagist
28GefährdetGesundheitsindex
roach-php/core
The complete web scraping toolkit for PHP.
PHP★ 1.455↓ 12K/Monat27. Juli 2026
Keine Lizenz27. Juli 2026 · Metriken 2.10.0
npm
20GefährdetGesundheitsindex
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0