Alle Tags
Katalog-Tag

#llm-inference

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

35 Einträge
Getaggt als „llm-inference“Geordnet nach Gesundheitsindex
Maven · PyPI
100AußergewöhnlichGesundheitsindex
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4K5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
PyPI · crates.io
99AußergewöhnlichGesundheitsindex
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.886↓ 59.4K/Monat28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
PyPI
98AußergewöhnlichGesundheitsindex
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6.25612. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI
97AußergewöhnlichGesundheitsindex
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda★ 6.263↓ 7.1M/Monat27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
Go · PyPI
97AußergewöhnlichGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.83728. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
96AußergewöhnlichGesundheitsindex
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8.782↓ 256.2K/Monat12. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI
95AußergewöhnlichGesundheitsindex
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5.422↓ 2.132/Monat2. Aug. 2026
Apache-2.02. Aug. 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 5.51028. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
monocle2ai/monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
Python★ 337↓ 50.8K/Monat28. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
93AußergewöhnlichGesundheitsindex
Lightning-AI/litgpt
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Python★ 13.6K↓ 15.5K/Monat27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
npm
93AußergewöhnlichGesundheitsindex
Nano-Collective/nanocoder
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
TypeScript★ 2.376↓ 9.541/Monat25. Aug. 2026
Eigene Lizenz25. Aug. 2026 · Metriken 2.10.0
npm · crates.io
93AußergewöhnlichGesundheitsindex
ruvnet/RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4.441↓ 432.2K/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
Packagist
92ExzellentGesundheitsindex
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2.012↓ 168.1K/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 2.10.0
Go · npm
89ExzellentGesundheitsindex
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617. Juli 2026
Apache-2.017. Juli 2026 · Metriken 2.10.0
Go · npm
87ExzellentGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
PyPI
86ExzellentGesundheitsindex
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 150↓ 1.822/Monat26. Aug. 2026
MIT26. Aug. 2026 · Metriken 2.10.0
PyPI
86ExzellentGesundheitsindex
algorithmicsuperintelligence/openevolve
Open-source implementation of AlphaEvolve
Python · JavaScript★ 7.27828. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2.506/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
npm
78GutGesundheitsindex
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 167↓ 6.355/Monat6. Sept. 2026
Eigene Lizenz6. Sept. 2026 · Metriken 2.10.0
npm
75GutGesundheitsindex
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/Monat22. Juli 2026
Apache-2.022. Juli 2026 · Metriken 2.10.0
PyPI
71GutGesundheitsindex
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3.360/Monat23. Aug. 2026
MIT23. Aug. 2026 · Metriken 2.10.0
Go · PyPI
65GutGesundheitsindex
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 2.10.0
PyPI
63MittelGesundheitsindex
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2.218/Monat15. Aug. 2026
Apache-2.015. Aug. 2026 · Metriken 2.10.0
Go
63MittelGesundheitsindex
flexigpt/inference-go
A single interface in Go to get inference from multiple llm/ai providers using their official SDKs
Go★ 23. Aug. 2026
MIT3. Aug. 2026 · Metriken 2.10.0
PyPI
63MittelGesundheitsindex
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516. Juli 2026
Eigene Lizenz16. Juli 2026 · Metriken 2.10.0
PyPI · crates.io
62MittelGesundheitsindex
theoddden/Terradev
An imperative command-line-interface for AI workload orchestration
Python★ 25↓ 3.845/Monat23. Aug. 2026
Apache-2.023. Aug. 2026 · Metriken 2.10.0
60MittelGesundheitsindex
eastriverlee/LLM.swift
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 86628. Juli 2026
MIT28. Juli 2026 · Metriken 2.10.0
60MittelGesundheitsindex
lean-dojo/LeanCopilot
LLMs as Copilots for Theorem Proving in Lean
C++★ 1.3074. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0
PyPI
59MittelGesundheitsindex
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
Go
59MittelGesundheitsindex
flexigpt/llmtools-go
LLM Tool implementations for Golang
Go★ 220. Aug. 2026
MIT20. Aug. 2026 · Metriken 2.10.0