全部标签
目录标签

#gguf

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

14 条记录
标签为“gguf”按健康指数排序
PyPI
87优秀健康指数
ggml-org/llama.cpp
LLM inference in C/C++
C++ · C★ 121.1K↓ 6.6M/月2026年7月21日
MIT2026年7月21日 · 指标 1.13.0
PyPI
85优秀健康指数
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日
Apache-2.02026年7月16日 · 指标 1.13.0
Go
74良好健康指数
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/月2026年7月14日
Apache-2.02026年7月14日 · 指标 1.13.0
PyPI
72良好健康指数
MakazhanAlpamys/Soup
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 742026年7月15日
Apache-2.02026年7月15日 · 指标 1.13.0
PyPI · crates.io · npm
71良好健康指数
alexsjones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
Go
65中等健康指数
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 5202026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
Go
62中等健康指数
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 276↓ 0/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
PyPI · crates.io
61中等健康指数
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7,559/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
Go
60中等健康指数
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 02026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
Go · PyPI
59中等健康指数
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 152026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
npm
59中等健康指数
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
Go
57中等健康指数
thxcode/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 2762026年7月15日
MIT2026年7月15日 · 指标 1.13.0
npm
55中等健康指数
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/月2026年7月16日
无许可证2026年7月16日 · 指标 1.13.0
PyPI
54中等健康指数
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 2142026年7月22日
无许可证2026年7月22日 · 指标 1.13.0