npm · crates.io · PyPI98Exceptionalhealth index
C++ · C# · Python★ 2,478↓ 137.8K/moJul 30, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7,796↓ 582.2K/moAug 28, 2026
npm95Exceptionalhealth index
TanStack/ai🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
TypeScript★ 2,907↓ 613.6K/moJul 22, 2026
npm94Exceptionalhealth index
TypeScript★ 272↓ 3.1M/moAug 27, 2026
PyPI · npm94Exceptionalhealth index

modelscope/FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
—94Exceptionalhealth index
tetherto/qvacQVAC - Local AI SDK and libraries for building private, cross-platform, peer-to-peer AI applications. Run LLMs, speech-to-text, translation, and more locally on Linux, macOS, Windows, Android, and iOS.
TypeScript · JavaScript · C++★ 321Jul 18, 2026
PyPI93Exceptionalhealth index
Python★ 459Aug 19, 2026
FluidInference/FluidAudioFrontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Swift★ 2,467Jul 16, 2026
npm90Excellenthealth index

AssemblyAI/assemblyai-node-sdkThe AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, audio intelligence models, as well as the latest LeMUR models.
TypeScript★ 75↓ 2.2M/moSep 6, 2026
crates.io89Excellenthealth index

screenpipe/screenpipeYC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps
Rust · TypeScript★ 20.7KAug 5, 2026
npm88Excellenthealth index
TypeScript★ 80↓ 169.2K/moJul 27, 2026
crates.io · npm87Excellenthealth index

cjpais/HandyA free, open source, and extensible speech-to-text application that works completely offline.
Rust · TypeScript★ 28.7KAug 5, 2026
PyPI · npm87Excellenthealth index

moonshine-ai/moonshineVery low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
C++ · C★ 10.9K↓ 49.9K/moAug 28, 2026
NuGet · npm87Excellenthealth index
C#★ 941Aug 31, 2026
PyPI86Excellenthealth index
Python★ 8,987↓ 9.2M/moAug 27, 2026
npm · crates.io86Excellenthealth index
drakulavich/kesha-voice-kitGive your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Rust · TypeScript★ 67↓ 1,029/moJul 28, 2026
crates.io · Maven · npm +184Excellenthealth index

k2-fsa/sherpa-onnxSpeech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
C++ · Python★ 14.4KAug 27, 2026
PyPI84Excellenthealth index
Python★ 204↓ 31.5K/moJul 19, 2026
C#★ 14Aug 25, 2026
PyPI83Excellenthealth index

modelscope/FunClipFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Python★ 6,201Aug 31, 2026
npm · crates.io83Excellenthealth index
silverstein/minutesEvery meeting, every idea, every voice note — searchable by your AI. Open-source, privacy-first conversation memory layer.
Rust★ 1,372↓ 5,342/moJul 23, 2026
PyPI81Excellenthealth index
istupakov/onnx-asrA lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/moJul 20, 2026
PyPI81Excellenthealth index

m-bain/whisperXWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.4K↓ 1.5M/moAug 5, 2026
TypeScript★ 2,080↓ 110.9K/moJul 21, 2026
GeiserX/whisper-subsJellyfin plugin for local AI-powered subtitle generation using Whisper - all processing stays on your server
C# · HTML★ 81Jul 24, 2026
TypeScript★ 47Jul 27, 2026
crates.io73Goodhealth index
altunenes/parakeet-rsvery fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Rust★ 378↓ 12.5K/moJul 29, 2026
dotdevdotdev/agentwire-devA self-hosted, keyboard-driven cockpit for running a whole fleet of Claude Code agents at once. Worktrees, command palette, scheduler, voice — every layer stacks, so you ship far more than hand-juggling terminals. Your machine, your keys, no telemetry.
Python · JavaScript★ 18↓ 4,499/moJul 19, 2026
C#★ 9Jul 21, 2026
gojargo/jargoA WebRTC-native, audio-first conversational-AI framework for Go.
Go★ 26Jul 17, 2026