Todas las etiquetas
Etiqueta del catálogo

#speech-to-text

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

55 registros
Con la etiqueta «speech-to-text»Ordenado por índice de salud
npm · crates.io · PyPI
98Excepcionalíndice de salud
microsoft/foundry-local
El repositorio no publica descripción.
C++ · C# · Python★ 2478↓ 137.8K/mes30 jul 2026
Licencia propia30 jul 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
Blaizzy/mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7796↓ 582.2K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
95Excepcionalíndice de salud
TanStack/ai
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
TypeScript★ 2907↓ 613.6K/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
npm
94Excepcionalíndice de salud
deepgram/deepgram-js-sdk
Official JavaScript SDK for Deepgram.
TypeScript★ 272↓ 3.1M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
94Excepcionalíndice de salud
tetherto/qvac
QVAC - Local AI SDK and libraries for building private, cross-platform, peer-to-peer AI applications. Run LLMs, speech-to-text, translation, and more locally on Linux, macOS, Windows, Android, and iOS.
TypeScript · JavaScript · C++★ 32118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
deepgram/deepgram-python-sdk
Official Python SDK for Deepgram.
Python★ 45919 ago 2026
MIT19 ago 2026 · métricas 2.10.0
92Excelenteíndice de salud
FluidInference/FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Swift★ 246716 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
npm
90Excelenteíndice de salud
AssemblyAI/assemblyai-node-sdk
The AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, audio intelligence models, as well as the latest LeMUR models.
TypeScript★ 75↓ 2.2M/mes6 sept 2026
MIT6 sept 2026 · métricas 2.10.0
crates.io
89Excelenteíndice de salud
screenpipe/screenpipe
YC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps
Rust · TypeScript★ 20.7K5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
npm
88Excelenteíndice de salud
HumeAI/hume-typescript-sdk
Add Hume AI to any TypeScript project
TypeScript★ 80↓ 169.2K/mes27 jul 2026
MIT27 jul 2026 · métricas 2.10.0
crates.io · npm
87Excelenteíndice de salud
cjpais/Handy
A free, open source, and extensible speech-to-text application that works completely offline.
Rust · TypeScript★ 28.7K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI · npm
87Excelenteíndice de salud
moonshine-ai/moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
C++ · C★ 10.9K↓ 49.9K/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
NuGet · npm
87Excelenteíndice de salud
sandrohanea/whisper.net
Whisper.net. Speech to text made simple using Whisper Models
C#★ 94131 ago 2026
MIT31 ago 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
Uberi/speech_recognition
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Python★ 8987↓ 9.2M/mes27 ago 2026
BSD-3-Clause27 ago 2026 · métricas 2.10.0
npm · crates.io
86Excelenteíndice de salud
drakulavich/kesha-voice-kit
Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Rust · TypeScript★ 67↓ 1029/mes28 jul 2026
MIT28 jul 2026 · métricas 2.10.0
crates.io · Maven · npm +1
84Excelenteíndice de salud
k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
C++ · Python★ 14.4K27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
84Excelenteíndice de salud
lfnovo/esperanto
A unified interface for various AI model providers
Python★ 204↓ 31.5K/mes19 jul 2026
MIT19 jul 2026 · métricas 2.10.0
84Excelenteíndice de salud
tryAGI/ElevenLabs
C# SDK for the ElevenLabs API -- text-to-speech and speech-to-text
C#★ 1425 ago 2026
MIT25 ago 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
modelscope/FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Python★ 620131 ago 2026
MIT31 ago 2026 · métricas 2.10.0
npm · crates.io
83Excelenteíndice de salud
silverstein/minutes
Every meeting, every idea, every voice note — searchable by your AI. Open-source, privacy-first conversation memory layer.
Rust★ 1372↓ 5342/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
istupakov/onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/mes20 jul 2026
MIT20 jul 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
m-bain/whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.4K↓ 1.5M/mes5 ago 2026
BSD-2-Clause5 ago 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
zerodytrash/TikTok-Live-Connector
Node.js library to receive live stream events (comments, gifts, etc.) in realtime from TikTok LIVE.
TypeScript★ 2080↓ 110.9K/mes21 jul 2026
AGPL-3.021 jul 2026 · métricas 2.10.0
NuGet
73Buenoíndice de salud
GeiserX/whisper-subs
Jellyfin plugin for local AI-powered subtitle generation using Whisper - all processing stays on your server
C# · HTML★ 8124 jul 2026
GPL-3.024 jul 2026 · métricas 2.10.0
npm
73Buenoíndice de salud
HumeAI/hume-react-sdk
Packages for using Hume AI and React
TypeScript★ 4727 jul 2026
Sin licencia27 jul 2026 · métricas 2.10.0
crates.io
73Buenoíndice de salud
altunenes/parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Rust★ 378↓ 12.5K/mes29 jul 2026
MIT29 jul 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
dotdevdotdev/agentwire-dev
A self-hosted, keyboard-driven cockpit for running a whole fleet of Claude Code agents at once. Worktrees, command palette, scheduler, voice — every layer stacks, so you ship far more than hand-juggling terminals. Your machine, your keys, no telemetry.
Python · JavaScript★ 18↓ 4499/mes19 jul 2026
Apache-2.019 jul 2026 · métricas 2.10.0
NuGet
73Buenoíndice de salud
fgilde/Nextended
Moved package nExt from CodePlex to github and updated to .net5
C#★ 921 jul 2026
GPL-3.021 jul 2026 · métricas 2.10.0
Go
73Buenoíndice de salud
gojargo/jargo
A WebRTC-native, audio-first conversational-AI framework for Go.
Go★ 2617 jul 2026
BSD-2-Clause17 jul 2026 · métricas 2.10.0