Усі теги
Тег каталогу

#llm-inference

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

19 записів
З тегом «llm-inference»Упорядковано за індексом здоров'я
Maven · PyPI
90Відміннийіндекс здоров'я
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.3K17 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
crates.io · PyPI
89Відміннийіндекс здоров'я
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7 52619 лип. 2026 р.
Власна ліцензія19 лип. 2026 р. · метрики 1.13.0
PyPI · crates.io
88Відміннийіндекс здоров'я
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7 51417 лип. 2026 р.
Власна ліцензія17 лип. 2026 р. · метрики 1.13.0
Go · PyPI
87Відміннийіндекс здоров'я
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5 70517 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
PyPI
85Відміннийіндекс здоров'я
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda · C++★ 5 988↓ 3M/міс21 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 1.13.0
Go · PyPI
85Відміннийіндекс здоров'я
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5 70517 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
PyPI
79Добрийіндекс здоров'я
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 4 96417 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
Packagist
78Добрийіндекс здоров'я
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2 012↓ 168.1K/міс15 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 1.13.0
Go · npm
75Добрийіндекс здоров'я
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
crates.io · npm
75Добрийіндекс здоров'я
ruvnet/ruvector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust★ 4 351↓ 0/міс13 лип. 2026 р.
MIT13 лип. 2026 р. · метрики 1.13.0
Go · npm
73Добрийіндекс здоров'я
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 1.13.0
PyPI
70Добрийіндекс здоров'я
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 14115 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 1.13.0
npm
65Помірнийіндекс здоров'я
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/міс22 лип. 2026 р.
Apache-2.022 лип. 2026 р. · метрики 1.13.0
npm
62Помірнийіндекс здоров'я
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 164↓ 11.5K/міс15 лип. 2026 р.
Власна ліцензія15 лип. 2026 р. · метрики 1.13.0
PyPI
60Помірнийіндекс здоров'я
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 лип. 2026 р.
Власна ліцензія16 лип. 2026 р. · метрики 1.13.0
Go · PyPI
59Помірнийіндекс здоров'я
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
PyPI
55Помірнийіндекс здоров'я
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 1.13.0
PyPI
55Помірнийіндекс здоров'я
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 1.13.0
npm
48У зоні ризикуіндекс здоров'я
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 16715 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 1.13.0