All tags
Catalogue tag

#speech-to-text

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

55 records
Tagged “speech-to-text”Ranked by health index
npm · crates.io · PyPI
98Exceptionalhealth index
microsoft/foundry-local
No repository description published.
C++ · C# · Python★ 2,478↓ 137.8K/moJul 30, 2026
Custom licenseJul 30, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
Blaizzy/mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7,796↓ 582.2K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm
95Exceptionalhealth index
TanStack/ai
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
TypeScript★ 2,907↓ 613.6K/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
npm
94Exceptionalhealth index
deepgram/deepgram-js-sdk
Official JavaScript SDK for Deepgram.
TypeScript★ 272↓ 3.1M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
94Exceptionalhealth index
tetherto/qvac
QVAC - Local AI SDK and libraries for building private, cross-platform, peer-to-peer AI applications. Run LLMs, speech-to-text, translation, and more locally on Linux, macOS, Windows, Android, and iOS.
TypeScript · JavaScript · C++★ 321Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
deepgram/deepgram-python-sdk
Official Python SDK for Deepgram.
Python★ 459Aug 19, 2026
MITAug 19, 2026 · metrics 2.10.0
92Excellenthealth index
FluidInference/FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Swift★ 2,467Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
npm
90Excellenthealth index
AssemblyAI/assemblyai-node-sdk
The AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, audio intelligence models, as well as the latest LeMUR models.
TypeScript★ 75↓ 2.2M/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
crates.io
89Excellenthealth index
screenpipe/screenpipe
YC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps
Rust · TypeScript★ 20.7KAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
npm
88Excellenthealth index
HumeAI/hume-typescript-sdk
Add Hume AI to any TypeScript project
TypeScript★ 80↓ 169.2K/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
crates.io · npm
87Excellenthealth index
cjpais/Handy
A free, open source, and extensible speech-to-text application that works completely offline.
Rust · TypeScript★ 28.7KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI · npm
87Excellenthealth index
moonshine-ai/moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
C++ · C★ 10.9K↓ 49.9K/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
NuGet · npm
87Excellenthealth index
sandrohanea/whisper.net
Whisper.net. Speech to text made simple using Whisper Models
C#★ 941Aug 31, 2026
MITAug 31, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
Uberi/speech_recognition
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Python★ 8,987↓ 9.2M/moAug 27, 2026
BSD-3-ClauseAug 27, 2026 · metrics 2.10.0
npm · crates.io
86Excellenthealth index
drakulavich/kesha-voice-kit
Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Rust · TypeScript★ 67↓ 1,029/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
crates.io · Maven · npm +1
84Excellenthealth index
k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
C++ · Python★ 14.4KAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
84Excellenthealth index
lfnovo/esperanto
A unified interface for various AI model providers
Python★ 204↓ 31.5K/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
84Excellenthealth index
tryAGI/ElevenLabs
C# SDK for the ElevenLabs API -- text-to-speech and speech-to-text
C#★ 14Aug 25, 2026
MITAug 25, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
modelscope/FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Python★ 6,201Aug 31, 2026
MITAug 31, 2026 · metrics 2.10.0
npm · crates.io
83Excellenthealth index
silverstein/minutes
Every meeting, every idea, every voice note — searchable by your AI. Open-source, privacy-first conversation memory layer.
Rust★ 1,372↓ 5,342/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
istupakov/onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/moJul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
m-bain/whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.4K↓ 1.5M/moAug 5, 2026
BSD-2-ClauseAug 5, 2026 · metrics 2.10.0
npm
78Goodhealth index
zerodytrash/TikTok-Live-Connector
Node.js library to receive live stream events (comments, gifts, etc.) in realtime from TikTok LIVE.
TypeScript★ 2,080↓ 110.9K/moJul 21, 2026
AGPL-3.0Jul 21, 2026 · metrics 2.10.0
NuGet
73Goodhealth index
GeiserX/whisper-subs
Jellyfin plugin for local AI-powered subtitle generation using Whisper - all processing stays on your server
C# · HTML★ 81Jul 24, 2026
GPL-3.0Jul 24, 2026 · metrics 2.10.0
npm
73Goodhealth index
HumeAI/hume-react-sdk
Packages for using Hume AI and React
TypeScript★ 47Jul 27, 2026
No licenseJul 27, 2026 · metrics 2.10.0
crates.io
73Goodhealth index
altunenes/parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Rust★ 378↓ 12.5K/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
dotdevdotdev/agentwire-dev
A self-hosted, keyboard-driven cockpit for running a whole fleet of Claude Code agents at once. Worktrees, command palette, scheduler, voice — every layer stacks, so you ship far more than hand-juggling terminals. Your machine, your keys, no telemetry.
Python · JavaScript★ 18↓ 4,499/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
NuGet
73Goodhealth index
fgilde/Nextended
Moved package nExt from CodePlex to github and updated to .net5
C#★ 9Jul 21, 2026
GPL-3.0Jul 21, 2026 · metrics 2.10.0
Go
73Goodhealth index
gojargo/jargo
A WebRTC-native, audio-first conversational-AI framework for Go.
Go★ 26Jul 17, 2026
BSD-2-ClauseJul 17, 2026 · metrics 2.10.0