All tags
Catalogue tag

#gguf

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

29 records
Tagged “gguf”Ranked by health index
PyPI
99Exceptionalhealth index
ggml-org/llama.cpp
LLM inference in C/C++
C++ · C★ 122.7KAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
npm
95Exceptionalhealth index
huggingface/huggingface.js
Use Hugging Face with JavaScript
TypeScript★ 2,503↓ 22.4M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go
89Excellenthealth index
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
crates.io · PyPI · npm
87Excellenthealth index
AlexsJones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2,251/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
npm
87Excellenthealth index
withcatai/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 2,162↓ 3.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
MakazhanAlpamys/Soup
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 74Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
PyPI · npm
86Excellenthealth index
n24q02m/mcp-core
Shared foundation for building MCP servers -- Streamable HTTP transport, OAuth 2.1, browser-based credential setup, and a shared embedding daemon.
Python · TypeScript★ 1↓ 17.7K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
crates.io · Maven
84Excellenthealth index
eugenehp/llama-cpp-rs
A wrapper around the llama-cpp library for rust, including new Sampler API from llama-cpp.
Rust★ 46↓ 7,448/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
Go
81Excellenthealth index
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 295Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
Go · npm
80Excellenthealth index
kdeps/kdeps
Run AI workflows locally. Or deploy them anywhere. AI agent framework in YAML — workflow pipelines + autonomous agent loop. NVIDIA Inception member. Build, deploy, export as Docker/K8s/ISO.
Go★ 35Jul 24, 2026
Apache-2.0Jul 24, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io
77Goodhealth index
ThreatFlux/gguf
A rust gguf library
Rust★ 6↓ 3,633/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go
75Goodhealth index
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 520Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
PyPI · crates.io
67Goodhealth index
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7,559/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
npm
67Goodhealth index
wundercorp/openmodel
Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
Apache-2.0Aug 7, 2026 · metrics 2.10.0
Go
65Goodhealth index
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 0Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go · PyPI
65Goodhealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI · crates.io
63Moderatehealth index
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0
npm
62Moderatehealth index
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
60Moderatehealth index
eastriverlee/LLM.swift
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 866Jul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
Maven · npm
60Moderatehealth index
integrallis/models
In-JVM small language model inference
Java★ 3Sep 4, 2026
Apache-2.0Sep 4, 2026 · metrics 2.10.0
npm
56Moderatehealth index
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026
No licenseJul 16, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
No licenseJul 22, 2026 · metrics 2.10.0
PyPI · npm · crates.io
53Moderatehealth index
JunHwan-Kwon/deepbom
Local static analysis and evidence generation for deployed AI model artifacts
JavaScript · Rust★ 2↓ 3,094/moSep 4, 2026
Apache-2.0Sep 4, 2026 · metrics 2.10.0
npm · crates.io · Maven
51Moderatehealth index
arusatech/llama-cpp-pro
Llama cpp + CapacitorJS support
Makefile · C · C++★ 7↓ 495/moJul 31, 2026
MITJul 31, 2026 · metrics 2.10.0
44Weakhealth index
calcuis/gguf-connector
gguf (GPT-Generated Unified Format) connector
Python★ 60Aug 31, 2026
MITAug 31, 2026 · metrics 2.10.0