PyPI99Exceptionalhealth index

huggingface/transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 163.3K↓ 179M/moAug 4, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7,796↓ 582.2K/moAug 28, 2026
npm94Exceptionalhealth index
TypeScript★ 272↓ 3.1M/moAug 27, 2026
PyPI · npm94Exceptionalhealth index

modelscope/FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
PyPI93Exceptionalhealth index
Python★ 459Aug 19, 2026
PyPI · npm92Excellenthealth index
Python · TypeScript · Swift★ 4,911↓ 773K/moAug 12, 2026
npm · Maven89Excellenthealth index
C++ · C · Metal★ 799↓ 43.6K/moAug 1, 2026
NuGet · npm87Excellenthealth index
C#★ 941Aug 31, 2026
PyPI86Excellenthealth index
Python★ 8,987↓ 9.2M/moAug 27, 2026
PyPI84Excellenthealth index
C★ 4,332Aug 12, 2026
PyPI84Excellenthealth index
Python★ 1,146↓ 1.4M/moAug 13, 2026
npm · PyPI · crates.io83Excellenthealth index
decibri/decibriCross-platform audio capture, playback, and voice activity detection for Python, Rust, and Node.js powered by a single Rust core.
Rust · Python · JavaScript★ 23↓ 19.3K/moJul 23, 2026
PyPI83Excellenthealth index

modelscope/FunClipFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Python★ 6,201Aug 31, 2026
PyPI81Excellenthealth index
istupakov/onnx-asrA lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/moJul 20, 2026
PyPI81Excellenthealth index

m-bain/whisperXWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.4K↓ 1.5M/moAug 5, 2026
crates.io73Goodhealth index
altunenes/parakeet-rsvery fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Rust★ 378↓ 12.5K/moJul 29, 2026
gojargo/jargoA WebRTC-native, audio-first conversational-AI framework for Go.
Go★ 26Jul 17, 2026
crates.io · PyPI69Goodhealth index
crispstrobe/crispasrC++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
C++ · Python · C★ 435Jul 16, 2026
mybigday/whisper.nodeAn another Node binding of whisper.cpp to make same API with whisper.rn as much as possible.
C++ · JavaScript · C★ 8↓ 43.9K/moJul 25, 2026
PyPI63Moderatehealth index
Python★ 24.7KAug 5, 2026
PyPI60Moderatehealth index
Python★ 3↓ 4,013/moAug 1, 2026
PyPI59Moderatehealth index

KoljaB/RealtimeSTTA robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
Python★ 10.1KAug 13, 2026
VoiceBlender/voiceblenderA programmable voice platform: SIP and WebRTC call control, multi-party mixing, recording, TTS/STT, and pluggable AI agents (ElevenLabs, VAPI, Pipecat, Deepgram) — all driven through a REST API, webhooks, and a WebSocket event stream
Go★ 97Jul 20, 2026
HenestrosaDev/audiotextA desktop application that transcribes audio from files, microphone input or YouTube videos with the option to translate the content and create subtitles.
Python★ 351Jul 18, 2026
JavaScript · C++★ 53Jul 18, 2026
C++★ 10.6KAug 12, 2026
PyPI33At Riskhealth index
Python · Shell★ 26Jul 22, 2026

alphacep/vosk-apiOffline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Jupyter Notebook★ 15K↓ 776.1K/moAug 12, 2026