crates.io · PyPI88优秀健康指数vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMsPython★ 86.1K↓ 0/月2026年7月13日Apache-2.02026年7月13日 · 指标 1.13.0
crates.io · PyPI82良好健康指数sgl-project/sglangSGLang is a high-performance serving framework for large language models and multimodal models.Python★ 30.3K↓ 274M/月2026年7月14日Apache-2.02026年7月14日 · 指标 1.13.0
PyPI · npm82良好健康指数unslothai/unslothUnsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.Python · TypeScript★ 68.5K↓ 2.3M/月2026年7月20日Apache-2.02026年7月20日 · 指标 1.13.0
Go81良好健康指数ollama/ollamaGet up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.Go · C★ 176.2K2026年7月15日MIT2026年7月15日 · 指标 1.13.0
Go80良好健康指数jmorganca/ollamaGet up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.Go · C★ 176.3K2026年7月17日MIT2026年7月17日 · 指标 1.13.0
PyPI · npm71良好健康指数lightseekorg/tokenspeedTokenSpeed is a speed-of-light LLM inference engine.Python★ 1,639↓ 1.7M/月2026年7月21日MIT2026年7月21日 · 指标 1.13.0
npm59中等健康指数therealtimex/node-llama-cppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation levelTypeScript★ 1↓ 41.2K/月2026年7月17日MIT2026年7月17日 · 指标 1.13.0