All tags
Catalogue tag

#llm-inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

35 records
Tagged “llm-inference”Ranked by health index
PyPI
57Moderatehealth index
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
53Moderatehealth index
ARahim3/mlx-dspark
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, Ornith-1.0, ternary Bonsai-27B.
Python · Swift★ 430↓ 5,564/moAug 19, 2026
MITAug 19, 2026 · metrics 2.10.0
npm
50Moderatehealth index
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 167Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
47Weakhealth index
MekayelAnik/vllm-cpu
Wheels & Docker images for running vLLM on CPU-only systems, optimized for different CPU instruction sets
Shell★ 8Aug 10, 2026
GPL-3.0Aug 10, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
openvinotoolkit/openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
C++★ 10.6KAug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0