Go · crates.io · npm71良好健康指数
teranos/QNTXQNTX = Experiential ꩜ Learning ⌬ System ≡ Attestation +
Go · TypeScript · Rust★ 32026年8月1日
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日
RubixML/MLA high-level machine learning and deep learning library for the PHP language.
PHP★ 2,199↓ 56.1K/月2026年7月22日
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/月2026年7月24日
Ericson246/npu-optimizeHardware-aware CLI that detects NPUs/GPUs, finds compatible GGUF models, and recommends optimal llama.cpp inference configs
Go★ 02026年7月17日
NevermindNilas/NeluxNeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3,141/月2026年7月19日

OpenCSGs/csgliteCSGLite is a lightweight local LLM inference platform. One command downloads, loads, and chats with models. It ships a web UI, OpenAI-compatible API, llama.cpp inference, resumable downloads, marketplace browsing, and one-click AI app and coding-agent setup—all in a single cross-platform binary.
Go · TypeScript★ 362026年8月22日

RDFLib/OWL-RLA simple implementation of the OWL2 RL Profile on top of RDFLib: it expands the graph with all possible triples that OWL RL defines. It can be used together with RDFLib to expand an RDFLib Graph object, or as a stand alone service with its own serialization.
HTML★ 1752026年8月8日

gvergnaud/ts-pattern🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
TypeScript★ 15.1K↓ 25.3M/月2026年8月12日

wundercorp/openmodelUse any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/月2026年8月7日
TypeScript★ 8↓ 2,051/月2026年7月23日
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 152026年7月17日
Python · Jupyter Notebook★ 662026年7月17日

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/月2026年8月20日

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/月2026年8月15日
Python★ 24.7K2026年8月5日
druide67/asiaiMulti-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/月2026年7月18日
jagmarques/nexusquantTraining-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 252026年7月16日

mcmcjs/mcmcjsCommand-line tools for Bayesian modelling, MCMC inference, and post-inference diagnostics across probabilistic programming languages.
TypeScript · Vue★ 8↓ 8,889/月2026年9月5日

seantalts/stanlistanli, the Stan Language Interpreter: op-graph executor over precompiled stan-math kernels. Compile and sample Stan models with no C++ toolchain.
C++ · Python★ 11↓ 7,415/月2026年8月30日
TypeScript★ 1↓ 6,389/月2026年7月17日

oramasearch/oramacoreOramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more utilities.
Rust★ 2602026年8月28日
Python · TypeScript★ 1↓ 2,921/月2026年8月3日

tangle-network/tcloudTypeScript SDK, CLI, agent, and relayer for Tangle AI Cloud — decentralized LLM inference
TypeScript★ 0↓ 36.9K/月2026年9月5日
ai-hypercomputer/jetstreamJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 4512026年7月18日
Python★ 1↓ 9,506/月2026年8月29日
Cuda · Python · C++★ 0↓ 2,509/月2026年9月1日
dshakes/firstpassRoute every LLM request to the cheapest model that provably passes your quality gate — with a signed, tamper-evident receipt for every decision. Proof over prediction.
Rust★ 2↓ 143/月2026年7月26日
google/jetstreamJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 4512026年7月18日
mohitsoni48/TurboLLMRun any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/月2026年7月16日