Todas las etiquetas
Etiqueta del catálogo

#inference

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

106 registros
Con la etiqueta «inference»Ordenado por índice de salud
Go · crates.io · npm
71Buenoíndice de salud
teranos/QNTX
QNTX = Experiential ꩜ Learning ⌬ System ≡ Attestation +
Go · TypeScript · Rust★ 31 ago 2026
MIT1 ago 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
Packagist
69Buenoíndice de salud
RubixML/ML
A high-level machine learning and deep learning library for the PHP language.
PHP★ 2199↓ 56.1K/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI · npm
69Buenoíndice de salud
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4850/mes24 jul 2026
Apache-2.024 jul 2026 · métricas 2.10.0
Go
67Buenoíndice de salud
Ericson246/npu-optimize
Hardware-aware CLI that detects NPUs/GPUs, finds compatible GGUF models, and recommends optimal llama.cpp inference configs
Go★ 017 jul 2026
MIT17 jul 2026 · métricas 2.10.0
PyPI
67Buenoíndice de salud
NevermindNilas/Nelux
NeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3141/mes19 jul 2026
AGPL-3.019 jul 2026 · métricas 2.10.0
Go · npm
67Buenoíndice de salud
OpenCSGs/csglite
CSGLite is a lightweight local LLM inference platform. One command downloads, loads, and chats with models. It ships a web UI, OpenAI-compatible API, llama.cpp inference, resumable downloads, marketplace browsing, and one-click AI app and coding-agent setup—all in a single cross-platform binary.
Go · TypeScript★ 3622 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
PyPI
67Buenoíndice de salud
RDFLib/OWL-RL
A simple implementation of the OWL2 RL Profile on top of RDFLib: it expands the graph with all possible triples that OWL RL defines. It can be used together with RDFLib to expand an RDFLib Graph object, or as a stand alone service with its own serialization.
HTML★ 1758 ago 2026
Licencia propia8 ago 2026 · métricas 2.10.0
npm
67Buenoíndice de salud
gvergnaud/ts-pattern
🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
TypeScript★ 15.1K↓ 25.3M/mes12 ago 2026
MIT12 ago 2026 · métricas 2.10.0
npm
67Buenoíndice de salud
wundercorp/openmodel
Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3140/mes7 ago 2026
Apache-2.07 ago 2026 · métricas 2.10.0
npm
65Buenoíndice de salud
amanharshx/ultralytics-mcp
MCP for Ultralytics Platform workflows, datasets, training, prediction, and model operations.
TypeScript★ 8↓ 2051/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
Go · PyPI
65Buenoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
65Buenoíndice de salud
skylight-org/sparse-attention-hub
Advancing the frontier of efficient AI
Python · Jupyter Notebook★ 6617 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI · crates.io
63Moderadoíndice de salud
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6494/mes20 ago 2026
MIT20 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.7K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
druide67/asiai
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4662/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
mcmcjs/mcmcjs
Command-line tools for Bayesian modelling, MCMC inference, and post-inference diagnostics across probabilistic programming languages.
TypeScript · Vue★ 8↓ 8889/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI · npm
63Moderadoíndice de salud
seantalts/stanli
stanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain.
C++ · Python★ 11↓ 7415/mes30 ago 2026
BSD-3-Clause30 ago 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
inference-sh/sdk-js
El repositorio no publica descripción.
TypeScript★ 1↓ 6389/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
crates.io · PyPI
62Moderadoíndice de salud
oramasearch/oramacore
OramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more utilities.
Rust★ 26028 ago 2026
AGPL-3.028 ago 2026 · métricas 2.10.0
PyPI · npm
60Moderadoíndice de salud
mauriciobenjamin700/ort-vision-sdk
El repositorio no publica descripción.
Python · TypeScript★ 1↓ 2921/mes3 ago 2026
MIT3 ago 2026 · métricas 2.10.0
npm
60Moderadoíndice de salud
tangle-network/tcloud
TypeScript SDK, CLI, agent, and relayer for Tangle AI Cloud — decentralized LLM inference
TypeScript★ 0↓ 36.9K/mes5 sept 2026
Sin licencia5 sept 2026 · métricas 2.10.0
PyPI
59Moderadoíndice de salud
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
59Moderadoíndice de salud
pictograph-io/pictograph-sdk
Complete python SDK, CLI, and API reference for integrating with the Pictograph.io platform.
Python★ 1↓ 9506/mes29 ago 2026
MIT29 ago 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
Hai-Wenxiang/fusedtok
Fused CUDA kernels for LLM inference - RMSNorm / RoPE / SwiGLU
Cuda · Python · C++★ 0↓ 2509/mes1 sept 2026
MIT1 sept 2026 · métricas 2.10.0
crates.io · PyPI
57Moderadoíndice de salud
dshakes/firstpass
Route every LLM request to the cheapest model that provably passes your quality gate — with a signed, tamper-evident receipt for every decision. Proof over prediction.
Rust★ 2↓ 143/mes26 jul 2026
Apache-2.026 jul 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
npm
56Moderadoíndice de salud
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6179/mes16 jul 2026
Sin licencia16 jul 2026 · métricas 2.10.0