All tags
Catalogue tag

#cuda

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

33 records
Tagged “cuda”Ranked by health index
crates.io · PyPI
88Excellenthealth index
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
PyPI
86Excellenthealth index
tenstorrent/tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
C++ · Python★ 1,584Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
85Excellenthealth index
cupy/cupy
NumPy & SciPy for GPU
Python · Cython★ 12.2K↓ 42.2K/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
PyPI
85Excellenthealth index
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda · C++★ 5,988↓ 3M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
PyPI
83Goodhealth index
NVIDIA/cccl
CUDA Core Compute Libraries
C++ · Cuda★ 2,420Jul 15, 2026
Custom licenseJul 15, 2026 · metrics 1.13.0
PyPI
83Goodhealth index
NVIDIA/warp
A Python framework for GPU-accelerated simulation, robotics, and machine learning.
Python · C++★ 6,871↓ 816.6K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
crates.io · PyPI
82Goodhealth index
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
82Goodhealth index
shader-slang/slang
Making it easier to work with shaders
C++ · Slang★ 5,439Jul 19, 2026
Custom licenseJul 19, 2026 · metrics 1.13.0
PyPI
80Goodhealth index
QMCPACK/qmcpack
Main repository for QMCPACK, an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support
C++ · Python★ 394Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
PyPI · npm
80Goodhealth index
msaad00/agent-bom
Open security scanner and self-hosted control plane for AI, MCP, and cloud. One evidence model — run scans in your environment, centralize findings, govern in your VPC.
Python · TypeScript★ 28↓ 5,301/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI
78Goodhealth index
NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
C++ · Cuda · Python★ 10.1K↓ 9,805/moJul 15, 2026
Custom licenseJul 15, 2026 · metrics 1.13.0
PyPI
78Goodhealth index
deepmodeling/abacus-develop
An electronic structure package based on either plane wave basis or numerical atomic orbitals.
C++★ 280Jul 20, 2026
LGPL-3.0Jul 20, 2026 · metrics 1.13.0
PyPI
76Goodhealth index
llnl/blt
A streamlined CMake build system foundation for developing HPC software
C++★ 294Jul 20, 2026
BSD-3-ClauseJul 20, 2026 · metrics 1.13.0
crates.io
75Goodhealth index
snipsco/tract
Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference
Rust★ 2,991↓ 912.8K/moJul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
PyPI
73Goodhealth index
NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python · C++★ 886Jul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
PyPI
72Goodhealth index
arborx/ArborX
Performance-portable geometric search library
C++★ 236Jul 21, 2026
BSD-3-ClauseJul 21, 2026 · metrics 1.13.0
PyPI
71Goodhealth index
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,577↓ 9.4M/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
crates.io · npm · Go +1
70Goodhealth index
ashvardanian/StringZilla
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops 🦖
C · C++ · Cuda★ 3,515↓ 2,257/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI · npm
70Goodhealth index
pikselkroken/pixlstash
PixlStash helps you find things in an image library that's gotten out of hand. It imports and tags your images automatically, then lets you search by content or face, sort the keepers from the rubbish, and serve data to other tools (like ComfyUI) with a REST API. Use the desktop version or run it headless as a server.
Python · Vue★ 76Jul 18, 2026
GPL-3.0Jul 18, 2026 · metrics 1.13.0
69Moderatehealth index
CliMA/RRTMGP.jl
Fast, GPU-ready atmospheric radiative transfer in Julia: the RTE solver with RRTMGP correlated-k gas optics.
Julia★ 64Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
RubyGems
66Moderatehealth index
sonots/cumo
Cumo (pronounced like "koomo") is CUDA aware numerical library whose interface is highly compatible with Ruby Numo
C · Ruby★ 99Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
Go
65Moderatehealth index
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 520Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 117Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
Go
61Moderatehealth index
eitamring/gocudrv
CUDA driver API in pure Go, no cgo: loads libcuda at runtime, embeds PTX, JIT-compiles through the driver
Go★ 16Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm · PyPI
60Moderatehealth index
invergent-ai/surogate
Training/Fine-tuning at the speed of light
C++ · Python · Cuda★ 806↓ 0/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
59Moderatehealth index
therealtimex/node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
crates.io
56Moderatehealth index
santhsecurity/vyre
Compiler-grade sequential GPU compute. Workgroup-local stacks, queues, hashmaps, dominator trees, fixed-point dataflow. CUDA + WGPU + SPIR-V with bit-exact conformance gate. Rust.
Rust★ 3↓ 30.6K/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI · crates.io
51Moderatehealth index
GuillaumeLessard/qector-decoder
Production-grade Rust/Python quantum error correction decoder. 25+ decoder families: MWPM Blossom, Union-Find, BP-OSD, LDPC/qLDPC, belief-matching, CUDA/GPU batch decoding, AutoDecoder 7-tier fallback. PyMatching/Stim/Sinter compatible. Ed25519 license verification. Benchmark evidence included.
Python★ 3Jul 22, 2026
Custom licenseJul 22, 2026 · metrics 1.13.0