All tags
Catalogue tag

#inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

106 records
Tagged “inference”Ranked by health index
Go · crates.io · npm
71Goodhealth index
teranos/QNTX
QNTX = Experiential ꩜ Learning ⌬ System ≡ Attestation +
Go · TypeScript · Rust★ 3Aug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
Packagist
69Goodhealth index
RubixML/ML
A high-level machine learning and deep learning library for the PHP language.
PHP★ 2,199↓ 56.1K/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI · npm
69Goodhealth index
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026
Apache-2.0Jul 24, 2026 · metrics 2.10.0
Go
67Goodhealth index
Ericson246/npu-optimize
Hardware-aware CLI that detects NPUs/GPUs, finds compatible GGUF models, and recommends optimal llama.cpp inference configs
Go★ 0Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
NevermindNilas/Nelux
NeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3,141/moJul 19, 2026
AGPL-3.0Jul 19, 2026 · metrics 2.10.0
Go · npm
67Goodhealth index
OpenCSGs/csglite
CSGLite is a lightweight local LLM inference platform. One command downloads, loads, and chats with models. It ships a web UI, OpenAI-compatible API, llama.cpp inference, resumable downloads, marketplace browsing, and one-click AI app and coding-agent setup—all in a single cross-platform binary.
Go · TypeScript★ 36Aug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
RDFLib/OWL-RL
A simple implementation of the OWL2 RL Profile on top of RDFLib: it expands the graph with all possible triples that OWL RL defines. It can be used together with RDFLib to expand an RDFLib Graph object, or as a stand alone service with its own serialization.
HTML★ 175Aug 8, 2026
Custom licenseAug 8, 2026 · metrics 2.10.0
npm
67Goodhealth index
gvergnaud/ts-pattern
🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
TypeScript★ 15.1K↓ 25.3M/moAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
npm
67Goodhealth index
wundercorp/openmodel
Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
Apache-2.0Aug 7, 2026 · metrics 2.10.0
npm
65Goodhealth index
amanharshx/ultralytics-mcp
MCP for Ultralytics Platform workflows, datasets, training, prediction, and model operations.
TypeScript★ 8↓ 2,051/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
Go · PyPI
65Goodhealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
65Goodhealth index
skylight-org/sparse-attention-hub
Advancing the frontier of efficient AI
Python · Jupyter Notebook★ 66Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI · crates.io
63Moderatehealth index
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.7KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
druide67/asiai
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mcmcjs/mcmcjs
Command-line tools for Bayesian modelling, MCMC inference, and post-inference diagnostics across probabilistic programming languages.
TypeScript · Vue★ 8↓ 8,889/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI · npm
63Moderatehealth index
seantalts/stanli
stanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain.
C++ · Python★ 11↓ 7,415/moAug 30, 2026
BSD-3-ClauseAug 30, 2026 · metrics 2.10.0
npm
62Moderatehealth index
inference-sh/sdk-js
No repository description published.
TypeScript★ 1↓ 6,389/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
crates.io · PyPI
62Moderatehealth index
oramasearch/oramacore
OramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more utilities.
Rust★ 260Aug 28, 2026
AGPL-3.0Aug 28, 2026 · metrics 2.10.0
PyPI · npm
60Moderatehealth index
mauriciobenjamin700/ort-vision-sdk
No repository description published.
Python · TypeScript★ 1↓ 2,921/moAug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
npm
60Moderatehealth index
tangle-network/tcloud
TypeScript SDK, CLI, agent, and relayer for Tangle AI Cloud — decentralized LLM inference
TypeScript★ 0↓ 36.9K/moSep 5, 2026
No licenseSep 5, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
pictograph-io/pictograph-sdk
Complete python SDK, CLI, and API reference for integrating with the Pictograph.io platform.
Python★ 1↓ 9,506/moAug 29, 2026
MITAug 29, 2026 · metrics 2.10.0
PyPI
57Moderatehealth index
Hai-Wenxiang/fusedtok
Fused CUDA kernels for LLM inference - RMSNorm / RoPE / SwiGLU
Cuda · Python · C++★ 0↓ 2,509/moSep 1, 2026
MITSep 1, 2026 · metrics 2.10.0
crates.io · PyPI
57Moderatehealth index
dshakes/firstpass
Route every LLM request to the cheapest model that provably passes your quality gate — with a signed, tamper-evident receipt for every decision. Proof over prediction.
Rust★ 2↓ 143/moJul 26, 2026
Apache-2.0Jul 26, 2026 · metrics 2.10.0
PyPI
57Moderatehealth index
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
npm
56Moderatehealth index
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026
No licenseJul 16, 2026 · metrics 2.10.0