All tags
Catalogue tag

#inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

54 records
Tagged “inference”Ranked by health index
PyPI
61Moderatehealth index
cozy-creator/python-gen-worker
A collection of worker runtimes (with Dockerfiles) + worker functions, to be used by the gen-orchestrator.
Python★ 0Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
Go
60Moderatehealth index
Ericson246/npu-optimize
Hardware-aware CLI that detects NPUs/GPUs, finds compatible GGUF models, and recommends optimal llama.cpp inference configs
Go★ 0Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
Go · PyPI
59Moderatehealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
59Moderatehealth index
druide67/asiai
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
59Moderatehealth index
skylight-org/sparse-attention-hub
Advancing the frontier of efficient AI
Python · Jupyter Notebook★ 66Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
57Moderatehealth index
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.4KJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
57Moderatehealth index
inference-sh/sdk-js
No repository description published.
TypeScript★ 1↓ 6,389/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
crates.io · PyPI
56Moderatehealth index
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust★ 31↓ 0/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
npm
55Moderatehealth index
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026
No licenseJul 16, 2026 · metrics 1.13.0
Go
53Moderatehealth index
wentbackward/llm-proxy
A superfast proxy and smart load-balancer for AI Inference — virtualize models, share local and provider back-ends, optimal caching, lock-in sampling parameters, debug message flow and obtain OTel metrics
Go★ 8Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm · crates.io · PyPI
52Moderatehealth index
Rust★ 10↓ 2,394/moJul 15, 2026
No licenseJul 15, 2026 · metrics 1.13.0
PyPI
52Moderatehealth index
ToPo-ToPo-ToPo/local-llm-server
ローカルLLM(mlx / mlx-vlm / llama.cpp / router)を OpenAI 互換 API として起動・管理する軽量サーバー
Python★ 0↓ 4,302/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
PyPI
52Moderatehealth index
h2non/filetype.py
Small, dependency-free, fast Python package to infer binary file types checking the magic numbers signature
Python★ 769Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Go · npm
52Moderatehealth index
opencsgs/csglite
CSGLite is a lightweight local LLM runner for the CSGHub platform. One command downloads, loads, and chats with models. It ships a web UI, OpenAI-compatible API, llama.cpp inference, resumable downloads, marketplace browsing, and one-click AI app and coding-agent setup—all in a single cross-platform binary.
Go★ 31↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
NuGet
50Moderatehealth index
mdesalvo/OWLSharp
Lightweight and friendly .NET library for realizing modern Semantic Web applications (OWL2, SWRL)
C#★ 16Jul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 1.13.0
Go
50Moderatehealth index
thalesfsp/inference
Provides building blocks to integrate with AI / LLM providers and built-in, common, providers.
Go★ 0Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Packagist · PyPI
49At riskhealth index
cognesy/instructor-php
Unified LLM API, structured data outputs with LLMs, and agent SDK - in PHP
PHP★ 325↓ 5,250/moJul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
PyPI
46At riskhealth index
mihir0209/AI_engine
Free AI inference router for developers - 21 providers, OpenAI-compatible
Python★ 4↓ 10.8K/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
Packagist
38At riskhealth index
JetBrains/phpstorm-stubs
PHP runtime & extensions header files for PhpStorm
PHP★ 1,390↓ 1.9M/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
Packagist
30At riskhealth index
RubixML/ML
A high-level machine learning and deep learning library for the PHP language.
PHP★ 2,199↓ 56.1K/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
27Criticalhealth index
declare-lab/reccon
This repository contains the dataset and the PyTorch implementations of the models from the paper Recognizing Emotion Cause in Conversations.
Python★ 190Jul 15, 2026
No licenseJul 15, 2026 · metrics 1.13.0