
roboflow/rf-detrRF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
Python★ 8,948↓ 497.6K/月2026年8月12日
PHP★ 1,181↓ 189.3K/月2026年8月22日
Python★ 49.3K↓ 1.4M/月2026年8月12日
Python★ 28↓ 17.5K/月2026年8月19日

XTLS/Xray-coreXray, Penetrates Everything. Also the best v2ray-core. Where the magic happens. An open platform for various uses.
Go★ 40.9K2026年8月5日
Python★ 171.5K↓ 13.7M/月2026年8月4日
C#★ 52026年7月22日
TypeScript · Swift · Kotlin★ 9,582↓ 2.7M/月2026年8月27日
Python★ 4,309↓ 210.7K/月2026年8月28日
Rust★ 2,462↓ 311K/月2026年7月24日
Go★ 1752026年9月5日
Python★ 15↓ 2,711/月2026年7月18日
Python★ 3692026年9月5日
Rust★ 38↓ 6,547/月2026年7月17日
TypeScript · JavaScript★ 108↓ 1,605/月2026年9月2日
cool-japan/oximediaOxiMedia is the Sovereign Media Framework - A patent-free, memory-safe multimedia processing library written in pure Rust. Pure Rust reconstruction of both FFmpeg (multimedia processing) and OpenCV (computer vision) — unified in a single cohesive framework.
Rust★ 224↓ 108/月2026年7月17日
sinameraji/kimiflarekimi k2.7 terminal based coding agent & harness running on your own Cloudflare account.
TypeScript★ 158↓ 3,221/月2026年7月22日
DustinTrap/kvm-pilotSmart hands for your AI agents — write-capable, multi-plane (KVM + BMC + SSH) MCP server for bare-metal control (PiKVM/GLKVM, Redfish/IPMI BMCs): gated, verified, audited. Beta: seeking hardware reports.
Python★ 0↓ 2,540/月2026年7月18日
JochenYang/luma-mcpMulti-Model Visual Understanding MCP Server, GLM-4.6V, DeepSeek-OCR (free), and Qwen3-VL-Flash. Provide visual processing capabilities for AI coding models that do not support image understanding.多模型视觉理解MCP服务器,GLM-4.6V、DeepSeek-OCR(免费)和Qwen3-VL-Flash等。为不支持图片理解的 AI 编码模型提供视觉处理能力。
TypeScript · JavaScript★ 83↓ 11.3K/月2026年7月26日
kenryu42/pi-grok-cliUse your X Premium or SuperGrok subscription in pi—with automatic vision routing for text-only models.
TypeScript★ 22↓ 2,291/月2026年7月23日
Rust★ 53↓ 3,841/月2026年8月9日
Rust★ 14↓ 2,690/月2026年8月8日

HUANGCHIHHUNGLeo/claude-real-videoLet Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Python★ 2,095↓ 6,592/月2026年8月31日

qirabot/qirabotAI vision GUI automation for browsers, mobile apps, desktops, and games — no DOM, no selectors. Standalone or on top of Playwright / Selenium / Appium / pytest.
Python★ 8↓ 5,010/月2026年8月6日
Python★ 1↓ 834/月2026年8月22日
qirabot/qirabot-pythonAI vision GUI automation for browsers, mobile apps, desktops, and games — no DOM, no selectors. Standalone or on top of Playwright / Selenium / Appium / pytest.
Python★ 6↓ 5,270/月2026年7月15日

strawlab/dltDLT (direct linear transform) algorithm for camera calibration
Rust★ 4↓ 2,112/月2026年8月8日

SIMON-WORLD/codex-deepseek-visionCodex vision bridge for DeepSeek V4 Flash: give text-only DeepSeek image capability in Codex. Local proxy turns pasted images and view_image into text via free GLM-4V-Flash or any OpenAI-compatible vision API. No GPU, no Ollama.
Python★ 72026年8月24日
getpipher/visionCapability-aware vision + paste extension for the pi coding agent. Delegates image analysis only when the active model is text-only; passes through natively for multimodal models (zero delegation).
TypeScript★ 6↓ 2,786/月2026年8月3日
TypeScript · JavaScript★ 1↓ 3,469/月2026年8月3日