全部标签
目录标签

#speculative-decoding

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

7 条记录
标签为“speculative-decoding”按健康指数排序
npm
87优秀健康指数
withcatai/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 2,162↓ 3.6M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
crates.io · npm · PyPI
86优秀健康指数
dphnAI/sonar
Large-scale LLM inference engine
C++ · Python★ 1,8332026年8月13日
AGPL-3.02026年8月13日 · 指标 2.10.0
PyPI · npm
81优秀健康指数
youssofal/MTPLX
3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Python · Swift★ 1,120↓ 2,694/月2026年8月2日
Apache-2.02026年8月2日 · 指标 2.10.0
PyPI · crates.io
63中等健康指数
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/月2026年8月20日
MIT2026年8月20日 · 指标 2.10.0
npm
62中等健康指数
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
53中等健康指数
ARahim3/mlx-dspark
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, Ornith-1.0, ternary Bonsai-27B.
Python · Swift★ 430↓ 5,564/月2026年8月19日
MIT2026年8月19日 · 指标 2.10.0
npm · crates.io · Maven
51中等健康指数
arusatech/llama-cpp-pro
Llama cpp + CapacitorJS support
Makefile · C · C++★ 7↓ 495/月2026年7月31日
MIT2026年7月31日 · 指标 2.10.0