Alle Tags
Katalog-Tag

#llm-inference

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

19 Einträge
Getaggt als „llm-inference“Geordnet nach Gesundheitsindex
Maven · PyPI
90ExzellentGesundheitsindex
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.3K17. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
crates.io · PyPI
89ExzellentGesundheitsindex
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.52619. Juli 2026
Eigene Lizenz19. Juli 2026 · Metriken 1.13.0
PyPI · crates.io
88ExzellentGesundheitsindex
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.51417. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 1.13.0
Go · PyPI
87ExzellentGesundheitsindex
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
PyPI
85ExzellentGesundheitsindex
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda · C++★ 5.988↓ 3M/Monat21. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0
Go · PyPI
85ExzellentGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
PyPI
79GutGesundheitsindex
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 4.96417. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
Packagist
78GutGesundheitsindex
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2.012↓ 168.1K/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
Go · npm
75GutGesundheitsindex
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
crates.io · npm
75GutGesundheitsindex
ruvnet/ruvector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust★ 4.351↓ 0/Monat13. Juli 2026
MIT13. Juli 2026 · Metriken 1.13.0
Go · npm
73GutGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0
PyPI
70GutGesundheitsindex
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 14115. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
npm
65MittelGesundheitsindex
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/Monat22. Juli 2026
Apache-2.022. Juli 2026 · Metriken 1.13.0
npm
62MittelGesundheitsindex
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 164↓ 11.5K/Monat15. Juli 2026
Eigene Lizenz15. Juli 2026 · Metriken 1.13.0
PyPI
60MittelGesundheitsindex
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516. Juli 2026
Eigene Lizenz16. Juli 2026 · Metriken 1.13.0
Go · PyPI
59MittelGesundheitsindex
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
PyPI
55MittelGesundheitsindex
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
PyPI
55MittelGesundheitsindex
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
npm
48GefährdetGesundheitsindex
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 16715. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0