Todas las etiquetas
Etiqueta del catálogo

#gguf

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

29 registros
Con la etiqueta «gguf»Ordenado por índice de salud
PyPI
99Excepcionalíndice de salud
ggml-org/llama.cpp
LLM inference in C/C++
C++ · C★ 122.7K4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
npm
95Excepcionalíndice de salud
huggingface/huggingface.js
Use Hugging Face with JavaScript
TypeScript★ 2503↓ 22.4M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go
89Excelenteíndice de salud
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2075 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
crates.io · PyPI · npm
87Excelenteíndice de salud
AlexsJones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2251/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
withcatai/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 2162↓ 3.6M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
MakazhanAlpamys/Soup
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 7415 jul 2026
Apache-2.015 jul 2026 · métricas 2.10.0
PyPI · npm
86Excelenteíndice de salud
n24q02m/mcp-core
Shared foundation for building MCP servers -- Streamable HTTP transport, OAuth 2.1, browser-based credential setup, and a shared embedding daemon.
Python · TypeScript★ 1↓ 17.7K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
crates.io · Maven
84Excelenteíndice de salud
eugenehp/llama-cpp-rs
A wrapper around the llama-cpp library for rust, including new Sampler API from llama-cpp.
Rust★ 46↓ 7448/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
Go
81Excelenteíndice de salud
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
Go★ 2955 sept 2026
MIT5 sept 2026 · métricas 2.10.0
Go · npm
80Excelenteíndice de salud
kdeps/kdeps
Run AI workflows locally. Or deploy them anywhere. AI agent framework in YAML — workflow pipelines + autonomous agent loop. NVIDIA Inception member. Build, deploy, export as Docker/K8s/ISO.
Go★ 3524 jul 2026
Apache-2.024 jul 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2506/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
crates.io
77Buenoíndice de salud
ThreatFlux/gguf
A rust gguf library
Rust★ 6↓ 3633/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go
75Buenoíndice de salud
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 52017 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3360/mes23 ago 2026
MIT23 ago 2026 · métricas 2.10.0
PyPI · crates.io
67Buenoíndice de salud
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7559/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
npm
67Buenoíndice de salud
wundercorp/openmodel
Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3140/mes7 ago 2026
Apache-2.07 ago 2026 · métricas 2.10.0
Go
65Buenoíndice de salud
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 017 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go · PyPI
65Buenoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI · crates.io
63Moderadoíndice de salud
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6494/mes20 ago 2026
MIT20 ago 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
60Moderadoíndice de salud
eastriverlee/LLM.swift
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 86628 jul 2026
MIT28 jul 2026 · métricas 2.10.0
Maven · npm
60Moderadoíndice de salud
integrallis/models
In-JVM small language model inference
Java★ 34 sept 2026
Apache-2.04 sept 2026 · métricas 2.10.0
npm
56Moderadoíndice de salud
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6179/mes16 jul 2026
Sin licencia16 jul 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 21422 jul 2026
Sin licencia22 jul 2026 · métricas 2.10.0
PyPI · npm · crates.io
53Moderadoíndice de salud
JunHwan-Kwon/deepbom
Local static analysis and evidence generation for deployed AI model artifacts
JavaScript · Rust★ 2↓ 3094/mes4 sept 2026
Apache-2.04 sept 2026 · métricas 2.10.0
npm · crates.io · Maven
51Moderadoíndice de salud
arusatech/llama-cpp-pro
Llama cpp + CapacitorJS support
Makefile · C · C++★ 7↓ 495/mes31 jul 2026
MIT31 jul 2026 · métricas 2.10.0
44Débilíndice de salud
calcuis/gguf-connector
gguf (GPT-Generated Unified Format) connector
Python★ 6031 ago 2026
MIT31 ago 2026 · métricas 2.10.0