Усі теги
Тег каталогу

#speech-to-text

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

31 запис
З тегом «speech-to-text»Упорядковано за індексом здоров'я
npm
82Добрийіндекс здоров'я
TanStack/ai
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
TypeScript★ 2 907↓ 613.6K/міс22 лип. 2026 р.
MIT22 лип. 2026 р. · метрики 1.13.0
npm
79Добрийіндекс здоров'я
deepgram/deepgram-js-sdk
Official JavaScript SDK for Deepgram.
TypeScript★ 268↓ 2.4M/міс19 лип. 2026 р.
MIT19 лип. 2026 р. · метрики 1.13.0
78Добрийіндекс здоров'я
FluidInference/FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Swift★ 2 46716 лип. 2026 р.
Apache-2.016 лип. 2026 р. · метрики 1.13.0
PyPI · npm
78Добрийіндекс здоров'я
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.4K21 лип. 2026 р.
MIT21 лип. 2026 р. · метрики 1.13.0
78Добрийіндекс здоров'я
tetherto/qvac
QVAC - Local AI SDK and libraries for building private, cross-platform, peer-to-peer AI applications. Run LLMs, speech-to-text, translation, and more locally on Linux, macOS, Windows, Android, and iOS.
TypeScript · JavaScript · C++★ 32118 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 1.13.0
npm · PyPI
77Добрийіндекс здоров'я
alibaba-damo-academy/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17 лип. 2026 р.
MIT17 лип. 2026 р. · метрики 1.13.0
npm · PyPI
77Добрийіндекс здоров'я
alibaba/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17 лип. 2026 р.
MIT17 лип. 2026 р. · метрики 1.13.0
npm
73Добрийіндекс здоров'я
AssemblyAI/assemblyai-node-sdk
The AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, audio intelligence models, as well as the latest LeMUR models.
TypeScript★ 76↓ 1.6M/міс15 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 1.13.0
PyPI
71Добрийіндекс здоров'я
lfnovo/esperanto
A unified interface for various AI model providers
Python★ 204↓ 31.5K/міс19 лип. 2026 р.
MIT19 лип. 2026 р. · метрики 1.13.0
PyPI
70Добрийіндекс здоров'я
m-bain/whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.2K21 лип. 2026 р.
BSD-2-Clause21 лип. 2026 р. · метрики 1.13.0
crates.io · Maven · npm +1
69Помірнийіндекс здоров'я
k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
C++ · Python★ 13.6K17 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
npm · crates.io
69Помірнийіндекс здоров'я
silverstein/minutes
Every meeting, every idea, every voice note — searchable by your AI. Open-source, privacy-first conversation memory layer.
Rust★ 1 372↓ 5 342/міс23 лип. 2026 р.
MIT23 лип. 2026 р. · метрики 1.13.0
npm
68Помірнийіндекс здоров'я
zerodytrash/TikTok-Live-Connector
Node.js library to receive live stream events (comments, gifts, etc.) in realtime from TikTok LIVE.
TypeScript★ 2 080↓ 110.9K/міс21 лип. 2026 р.
AGPL-3.021 лип. 2026 р. · метрики 1.13.0
PyPI
64Помірнийіндекс здоров'я
dotdevdotdev/agentwire-dev
A self-hosted, keyboard-driven cockpit for running a whole fleet of Claude Code agents at once. Worktrees, command palette, scheduler, voice — every layer stacks, so you ship far more than hand-juggling terminals. Your machine, your keys, no telemetry.
Python · JavaScript★ 18↓ 4 499/міс19 лип. 2026 р.
Apache-2.019 лип. 2026 р. · метрики 1.13.0
NuGet
64Помірнийіндекс здоров'я
fgilde/Nextended
Moved package nExt from CodePlex to github and updated to .net5
C#★ 921 лип. 2026 р.
GPL-3.021 лип. 2026 р. · метрики 1.13.0
Go
63Помірнийіндекс здоров'я
gojargo/jargo
A WebRTC-native, audio-first conversational-AI framework for Go.
Go★ 2617 лип. 2026 р.
BSD-2-Clause17 лип. 2026 р. · метрики 1.13.0
crates.io · PyPI
61Помірнийіндекс здоров'я
crispstrobe/crispasr
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
C++ · Python · C★ 43516 лип. 2026 р.
MIT16 лип. 2026 р. · метрики 1.13.0
PyPI · crates.io
61Помірнийіндекс здоров'я
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7 559/міс17 лип. 2026 р.
MIT17 лип. 2026 р. · метрики 1.13.0
npm · PyPI
59Помірнийіндекс здоров'я
arcforgelabs/dictate
Desktop dictation that types into the focused app. Configurable push-to-talk, local/API transcription, tray settings, and recent history.
Python · JavaScript★ 0↓ 8 233/міс16 лип. 2026 р.
MIT16 лип. 2026 р. · метрики 1.13.0
npm · crates.io
59Помірнийіндекс здоров'я
snomiao/otoji
realtime speech ⇄ text, 音を字に — wire mic → STT → translate → speech as an on-device voice graph that spans your devices over WebRTC. Runs in the browser (transformers.js/ONNX/WebGPU); no API keys, nothing leaves the device by default.
TypeScript · Rust★ 1↓ 4 782/міс15 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 1.13.0
PyPI
57Помірнийіндекс здоров'я
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.4K18 лип. 2026 р.
MIT18 лип. 2026 р. · метрики 1.13.0
Go
55Помірнийіндекс здоров'я
VoiceBlender/voiceblender
A programmable voice platform: SIP and WebRTC call control, multi-party mixing, recording, TTS/STT, and pluggable AI agents (ElevenLabs, VAPI, Pipecat, Deepgram) — all driven through a REST API, webhooks, and a WebSocket event stream
Go★ 9720 лип. 2026 р.
MIT20 лип. 2026 р. · метрики 1.13.0
npm
52Помірнийіндекс здоров'я
arach/vox
Local-first macOS transcription runtime with Swift services, a Bun CLI, and a TypeScript SDK.
Swift · TypeScript★ 519 лип. 2026 р.
Без ліцензії19 лип. 2026 р. · метрики 1.13.0
npm · RubyGems · Go +2
50Помірнийіндекс здоров'я
runapi-ai/elevenlabs-sdk
RunAPI ElevenLabs SDK for text-to-speech, dialogue generation, sound effects, speech transcription, and audio isolation workflows in JavaScript, Python, Ruby, Go, Java, and PHP
Java★ 0↓ 254/міс21 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 1.13.0
PyPI
49У зоні ризикуіндекс здоров'я
Leon-Sander/Local-Multimodal-AI-Chat
Self-hostable multimodal chat with local LLMs (Ollama/OpenAI): PDF RAG, image chat, and Whisper voice, Streamlit + Docker.
Python★ 20318 лип. 2026 р.
GPL-3.018 лип. 2026 р. · метрики 1.13.0
PyPI
44У зоні ризикуіндекс здоров'я
HenestrosaDev/audiotext
A desktop application that transcribes audio from files, microphone input or YouTube videos with the option to translate the content and create subtitles.
Python★ 35118 лип. 2026 р.
Власна ліцензія18 лип. 2026 р. · метрики 1.13.0
npm
41У зоні ризикуіндекс здоров'я
TranscribeJs/transcribe.js
Monorepo for Transcribe.js
JavaScript · C++★ 5318 лип. 2026 р.
MIT18 лип. 2026 р. · метрики 1.13.0
PyPI
39У зоні ризикуіндекс здоров'я
collectivat/cmusphinx-models
Acoustic and language models for minorised languages.
Python · Shell★ 2622 лип. 2026 р.
AGPL-3.022 лип. 2026 р. · метрики 1.13.0
crates.io · npm
35У зоні ризикуіндекс здоров'я
cjpais/Handy
A free, open source, and extensible speech-to-text application that works completely offline.
Rust · TypeScript★ 27K20 лип. 2026 р.
MIT20 лип. 2026 р. · метрики 1.13.0
PyPI
33У зоні ризикуіндекс здоров'я
istupakov/onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/міс20 лип. 2026 р.
MIT20 лип. 2026 р. · метрики 1.13.0