All tags
Catalogue tag

#gguf

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

14 records
Tagged “gguf”Ranked by health index
PyPI
87Excellenthealth index
ggml-org/llama.cpp
LLM inference in C/C++
C++ · C★ 121.1K↓ 6.6M/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
PyPI
85Excellenthealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
Go
74Goodhealth index
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
PyPI
72Goodhealth index
MakazhanAlpamys/Soup
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 74Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
PyPI · crates.io · npm
71Goodhealth index
alexsjones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Go
65Moderatehealth index
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 520Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go
62Moderatehealth index
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 276↓ 0/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
PyPI · crates.io
61Moderatehealth index
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7,559/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Go
60Moderatehealth index
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 0Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
Go · PyPI
59Moderatehealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
npm
59Moderatehealth index
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Go
57Moderatehealth index
thxcode/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 276Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
npm
55Moderatehealth index
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026
No licenseJul 16, 2026 · metrics 1.13.0
PyPI
54Moderatehealth index
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
No licenseJul 22, 2026 · metrics 1.13.0