Go · crates.io · npm71Goodhealth index
teranos/QNTXQNTX = Experiential ꩜ Learning ⌬ System ≡ Attestation +
Go · TypeScript · Rust★ 3Aug 1, 2026
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
Packagist69Goodhealth index
RubixML/MLA high-level machine learning and deep learning library for the PHP language.
PHP★ 2,199↓ 56.1K/moJul 22, 2026
PyPI · npm69Goodhealth index
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026
Ericson246/npu-optimizeHardware-aware CLI that detects NPUs/GPUs, finds compatible GGUF models, and recommends optimal llama.cpp inference configs
Go★ 0Jul 17, 2026
NevermindNilas/NeluxNeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3,141/moJul 19, 2026
Go · npm67Goodhealth index

OpenCSGs/csgliteCSGLite is a lightweight local LLM inference platform. One command downloads, loads, and chats with models. It ships a web UI, OpenAI-compatible API, llama.cpp inference, resumable downloads, marketplace browsing, and one-click AI app and coding-agent setup—all in a single cross-platform binary.
Go · TypeScript★ 36Aug 22, 2026

RDFLib/OWL-RLA simple implementation of the OWL2 RL Profile on top of RDFLib: it expands the graph with all possible triples that OWL RL defines. It can be used together with RDFLib to expand an RDFLib Graph object, or as a stand alone service with its own serialization.
HTML★ 175Aug 8, 2026

gvergnaud/ts-pattern🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
TypeScript★ 15.1K↓ 25.3M/moAug 12, 2026

wundercorp/openmodelUse any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
TypeScript★ 8↓ 2,051/moJul 23, 2026
Go · PyPI65Goodhealth index
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Python · Jupyter Notebook★ 66Jul 17, 2026
PyPI · crates.io63Moderatehealth index

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
PyPI63Moderatehealth index

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
PyPI63Moderatehealth index
Python★ 24.7KAug 5, 2026
PyPI63Moderatehealth index
druide67/asiaiMulti-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
PyPI63Moderatehealth index
jagmarques/nexusquantTraining-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
npm63Moderatehealth index

mcmcjs/mcmcjsCommand-line tools for Bayesian modelling, MCMC inference, and post-inference diagnostics across probabilistic programming languages.
TypeScript · Vue★ 8↓ 8,889/moSep 5, 2026
PyPI · npm63Moderatehealth index

seantalts/stanlistanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain.
C++ · Python★ 11↓ 7,415/moAug 30, 2026
npm62Moderatehealth index
TypeScript★ 1↓ 6,389/moJul 17, 2026
crates.io · PyPI62Moderatehealth index

oramasearch/oramacoreOramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more utilities.
Rust★ 260Aug 28, 2026
PyPI · npm60Moderatehealth index
Python · TypeScript★ 1↓ 2,921/moAug 3, 2026
npm60Moderatehealth index

tangle-network/tcloudTypeScript SDK, CLI, agent, and relayer for Tangle AI Cloud — decentralized LLM inference
TypeScript★ 0↓ 36.9K/moSep 5, 2026
PyPI59Moderatehealth index
ai-hypercomputer/jetstreamJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
PyPI59Moderatehealth index
Python★ 1↓ 9,506/moAug 29, 2026
PyPI57Moderatehealth index
Cuda · Python · C++★ 0↓ 2,509/moSep 1, 2026
crates.io · PyPI57Moderatehealth index
dshakes/firstpassRoute every LLM request to the cheapest model that provably passes your quality gate — with a signed, tamper-evident receipt for every decision. Proof over prediction.
Rust★ 2↓ 143/moJul 26, 2026
PyPI57Moderatehealth index
google/jetstreamJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
npm56Moderatehealth index
mohitsoni48/TurboLLMRun any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026