Alle Tags
Katalog-Tag

#crawler

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

85 Einträge
Getaggt als „crawler“Geordnet nach Gesundheitsindex
Packagist
57MittelGesundheitsindex
JayBizzle/Laravel-Crawler-Detect
A Laravel wrapper for CrawlerDetect - the web crawler detection library
PHP★ 324↓ 61.2K/Monat28. Juli 2026
MIT28. Juli 2026 · Metriken 2.10.0
PyPI
57MittelGesundheitsindex
codelucas/newspaper
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/Monat12. Aug. 2026
MIT12. Aug. 2026 · Metriken 2.10.0
Go
56MittelGesundheitsindex
x-way/crawlerdetect
Golang module to detect bots and crawlers via the user agent
Go★ 7017. Juli 2026
MIT17. Juli 2026 · Metriken 2.10.0
npm
54MittelGesundheitsindex
codepurse/SEOCORE
Enterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health audits.
TypeScript★ 10917. Juli 2026
MIT17. Juli 2026 · Metriken 2.10.0
PyPI
54MittelGesundheitsindex
tobybgy-lsd/web-agent-runtime-bench
Local-first failure diagnosis, auto collection, repair planning, AI handoff, and verification for Playwright, crawler, RPA, and agent workflows.
Python · JavaScript★ 1↓ 3.325/Monat26. Juli 2026
Eigene Lizenz26. Juli 2026 · Metriken 2.10.0
npm
51MittelGesundheitsindex
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/Monat22. Juli 2026
Keine Lizenz22. Juli 2026 · Metriken 2.10.0
Packagist · npm
50MittelGesundheitsindex
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/Monat29. Juli 2026
MIT29. Juli 2026 · Metriken 2.10.0
Go
50MittelGesundheitsindex
mjc-gh/virgo
Convert a webpage into plaintext or markdown using the Chrome DevTool Protocol
Go · JavaScript★ 117. Juli 2026
BSD-3-Clause17. Juli 2026 · Metriken 2.10.0
Go
50MittelGesundheitsindex
quantmind-br/repodocs
Go CLI for extracting websites, repositories, sitemaps, and package docs into structured Markdown
Go★ 025. Aug. 2026
MIT25. Aug. 2026 · Metriken 2.10.0
crates.io
50MittelGesundheitsindex
spider-rs/spider_firewall
Firewall for Rust
Rust★ 1↓ 2.668/Monat16. Juli 2026
MIT16. Juli 2026 · Metriken 2.10.0
npm · Go · PyPI
48SchwachGesundheitsindex
AlphaTechini/doc-fetch
Dynamic documentation fetching CLI that converts entire documentation sites to single markdown files for AI/LLM consumption
Svelte · Go★ 1↓ 2.547/Monat27. Juli 2026
Keine Lizenz27. Juli 2026 · Metriken 2.10.0
Packagist
42SchwachGesundheitsindex
crawlbase/crawlbase-php
A lightweight, dependency free PHP class that acts as wrapper for Crawlbase API
PHP★ 16↓ 2.779/Monat15. Juli 2026
Apache-2.015. Juli 2026 · Metriken 2.10.0
Packagist
41SchwachGesundheitsindex
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2.07827. Juli 2026
Eigene Lizenz27. Juli 2026 · Metriken 2.10.0
crates.io
39SchwachGesundheitsindex
zekroTJA/r34-crawler
A simple CLI tool to fetch and download images from rule34.xxx
Rust★ 422. Juli 2026
MIT22. Juli 2026 · Metriken 2.10.0
npm
36SchwachGesundheitsindex
pavlealeksic/playwright-afp
Stop website fingerprinting techniques playwright edition
JavaScript★ 19↓ 4.152/Monat2. Sept. 2026
MIT2. Sept. 2026 · Metriken 2.10.0
NuGet
35SchwachGesundheitsindex
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 6231. Juli 2026
Eigene Lizenz31. Juli 2026 · Metriken 2.10.0
PyPI
34GefährdetGesundheitsindex
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8K12. Aug. 2026
BSD-3-Clause12. Aug. 2026 · Metriken 2.10.0
npm
30GefährdetGesundheitsindex
hardbulls/wbsc-crawler
Keine Repository-Beschreibung veröffentlicht.
TypeScript★ 1↓ 2.234/Monat25. Juli 2026
Keine Lizenz25. Juli 2026 · Metriken 2.10.0
Maven
29GefährdetGesundheitsindex
code4craft/webmagic
A scalable web crawler framework for Java.
Java · HTML★ 11.7K27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
npm
29GefährdetGesundheitsindex
crawlbase/crawlbase-node
Fast dependency free library for Crawlbase API
JavaScript★ 919. Juli 2026
Apache-2.019. Juli 2026 · Metriken 2.10.0
Go
28GefährdetGesundheitsindex
Synoppy/synoppy-go
Official Go SDK for Synoppy — the web-data layer for AI agents. Read, crawl, map, extract, classify & enrich any website on one key. Standard library only.
Go★ 128. Juli 2026
MIT28. Juli 2026 · Metriken 2.10.0
Maven
23GefährdetGesundheitsindex
ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。
Java★ 11.4K27. Aug. 2026
MIT27. Aug. 2026 · Metriken 2.10.0
PyPI
22GefährdetGesundheitsindex
0-EternalJunior-0/GraphCrawler
Python бібліотека для сканування веб-сайтів та побудови графу їх структури.
Python · HTML★ 1↓ 4.246/Monat1. Aug. 2026
MIT1. Aug. 2026 · Metriken 2.10.0
npm · PyPI
20GefährdetGesundheitsindex
88899/gitmen-lottery
彩票 数据 双色球 大乐透 快开 超级大乐透 预测 仅供学习
Python · JavaScript★ 927. Juli 2026
Keine Lizenz27. Juli 2026 · Metriken 2.10.0
npm
20GefährdetGesundheitsindex
karthikuj/sasori
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0