Todas las etiquetas
Etiqueta del catálogo

#gguf

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

14 registros
Con la etiqueta «gguf»Ordenado por índice de salud
PyPI
87Excelenteíndice de salud
ggml-org/llama.cpp
LLM inference in C/C++
C++ · C★ 121.1K↓ 6.6M/mes21 jul 2026
MIT21 jul 2026 · métricas 1.13.0
PyPI
85Excelenteíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
Go
74Buenoíndice de salud
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI
72Buenoíndice de salud
MakazhanAlpamys/Soup
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 7415 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
PyPI · crates.io · npm
71Buenoíndice de salud
alexsjones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
Go
65Moderadoíndice de salud
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 52017 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go
62Moderadoíndice de salud
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 276↓ 0/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
PyPI · crates.io
61Moderadoíndice de salud
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7559/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
Go
60Moderadoíndice de salud
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 017 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go · PyPI
59Moderadoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
npm
59Moderadoíndice de salud
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
Go
57Moderadoíndice de salud
thxcode/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 27615 jul 2026
MIT15 jul 2026 · métricas 1.13.0
npm
55Moderadoíndice de salud
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6179/mes16 jul 2026
Sin licencia16 jul 2026 · métricas 1.13.0
PyPI
54Moderadoíndice de salud
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 21422 jul 2026
Sin licencia22 jul 2026 · métricas 1.13.0