PyPI88Excellenthealth index
NVIDIA/cudnn-frontendcuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python · C++★ 886Jul 21, 2026
npm · PyPI88Excellenthealth index

roflcoopter/viseronSelf-hosted, local only NVR and AI Computer Vision software. With features such as object detection, motion detection, face recognition and more, it gives you the power to keep an eye on your home, office or any other place you want to monitor.
Python · TypeScript★ 3,464Sep 2, 2026
PyPI87Excellenthealth index
arbor-sim/arborThe Arbor multi-compartment neural network simulation library.
C++ · AGS Script★ 136Jul 25, 2026
npm87Excellenthealth index

withcatai/node-llama-cppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 2,162↓ 3.6M/moAug 27, 2026
PyPI86Excellenthealth index
C++ · Python★ 4,645↓ 12.4M/moAug 27, 2026
PyPI86Excellenthealth index
C++★ 236Jul 21, 2026
crates.io · npm · PyPI86Excellenthealth index
C++ · Python★ 1,833Aug 13, 2026
PyPI84Excellenthealth index

NVIDIA/ncclOptimized primitives for collective multi-GPU communication
C++ · Cuda · C★ 4,988↓ 634.8K/moAug 12, 2026
PyPI84Excellenthealth index
C++★ 13.9KAug 12, 2026
PyPI · npm83Excellenthealth index
pikselkroken/pixlstashPixlStash helps you find things in an image library that's gotten out of hand. It imports and tags your images automatically, then lets you search by content or face, sort the keepers from the rubbish, and serve data to other tools (like ComfyUI) with a REST API. Use the desktop version or run it headless as a server.
Python · Vue★ 76Jul 18, 2026
crates.io83Excellenthealth index
Rust★ 105↓ 217.6K/moJul 29, 2026
CliMA/RRTMGP.jlFast, GPU-ready atmospheric radiative transfer in Julia: the RTE solver with RRTMGP correlated-k gas optics.
Julia★ 64Jul 17, 2026
crates.io · npm · Go +181Excellenthealth index
ashvardanian/StringZillaUp to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops 🦖
C · C++ · Cuda★ 3,515↓ 2,257/moJul 21, 2026
PyPI80Excellenthealth index

XuehaiPan/nvitopAn interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management.
Python★ 7,100Aug 12, 2026
C++★ 1,313Jul 27, 2026
PyPI · crates.io77Goodhealth index

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
crates.io77Goodhealth index
Rust★ 1,206↓ 939.6K/moAug 16, 2026
crates.io77Goodhealth index
Rust · Cuda★ 5↓ 9,587/moAug 22, 2026
hybridgroup/yzmaGo with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 520Jul 17, 2026
npm · crates.io75Goodhealth index
Rust · TypeScript★ 152↓ 2,518/moAug 3, 2026
RubyGems75Goodhealth index
sonots/cumoCumo (pronounced like "koomo") is CUDA aware numerical library whose interface is highly compatible with Ruby Numo
C · Ruby★ 99Jul 17, 2026
Python · Cuda★ 117Jul 18, 2026
C++ · Python · Cuda★ 806↓ 0/moJul 21, 2026

ooples/AiDotNet.TensorsThe fastest .NET tensor library. Beats MathNet (6x), NumSharp (3200x), matches TorchSharp CPU - pure managed C# with hand-tuned AVX2/FMA SIMD kernels. Optional CUDA/OpenCL GPU acceleration.
C#★ 11Sep 6, 2026
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
crates.io69Goodhealth index
santhreal/vyreCompiler-grade sequential GPU compute. Workgroup-local stacks, queues, hashmaps, dominator trees, fixed-point dataflow. CUDA + WGPU + SPIR-V with bit-exact conformance gate. Rust.
Rust★ 3↓ 67.2K/moAug 1, 2026
PyPI · crates.io67Goodhealth index

Egoist-Machines/LodeDBWorld's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated
Python · Rust★ 90Aug 5, 2026
eitamring/gocudrvCUDA driver API in pure Go, no cgo: loads libcuda at runtime, embeds PTX, JIT-compiles through the driver
Go★ 16Jul 18, 2026
Python★ 25↓ 43.4K/moAug 24, 2026
mybigday/whisper.nodeAn another Node binding of whisper.cpp to make same API with whisper.rn as much as possible.
C++ · JavaScript · C★ 8↓ 43.9K/moJul 25, 2026