全部标签
目录标签

#asr

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

27 条记录
标签为“asr”按健康指数排序
PyPI
99卓越健康指数
NVIDIA-NeMo/Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Python · Jupyter Notebook★ 18.1K2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
npm
94卓越健康指数
deepgram/deepgram-js-sdk
Official JavaScript SDK for Deepgram.
TypeScript★ 272↓ 3.1M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI
93卓越健康指数
deepgram/deepgram-python-sdk
Official Python SDK for Deepgram.
Python★ 4592026年8月19日
MIT2026年8月19日 · 指标 2.10.0
92优秀健康指数
FluidInference/FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Swift★ 2,4672026年7月16日
Apache-2.02026年7月16日 · 指标 2.10.0
npm
90优秀健康指数
AssemblyAI/assemblyai-node-sdk
The AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, audio intelligence models, as well as the latest LeMUR models.
TypeScript★ 75↓ 2.2M/月2026年9月6日
MIT2026年9月6日 · 指标 2.10.0
PyPI
87优秀健康指数
mbailey/voicemode
Natural voice conversations with Claude Code
Python★ 1,308↓ 22.8K/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI · npm
87优秀健康指数
moonshine-ai/moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
C++ · C★ 10.9K↓ 49.9K/月2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
npm · crates.io
86优秀健康指数
drakulavich/kesha-voice-kit
Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Rust · TypeScript★ 67↓ 1,029/月2026年7月28日
MIT2026年7月28日 · 指标 2.10.0
PyPI
84优秀健康指数
cmusphinx/pocketsphinx
A small speech recognizer
C★ 4,3322026年8月12日
自定义许可证2026年8月12日 · 指标 2.10.0
crates.io · Maven · npm +1
84优秀健康指数
k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
C++ · Python★ 14.4K2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
PyPI
83优秀健康指数
modelscope/FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Python★ 6,2012026年8月31日
MIT2026年8月31日 · 指标 2.10.0
PyPI
81优秀健康指数
istupakov/onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
Python★ 347↓ 244.7K/月2026年7月20日
MIT2026年7月20日 · 指标 2.10.0
PyPI
81优秀健康指数
m-bain/whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python★ 23.4K↓ 1.5M/月2026年8月5日
BSD-2-Clause2026年8月5日 · 指标 2.10.0
PyPI · npm
78良好健康指数
bengizmo/voxint
该仓库未发布描述。
Python★ 3↓ 2,239/月2026年8月21日
Apache-2.02026年8月21日 · 指标 2.10.0
npm
78良好健康指数
speechmatics/speechmatics-js-sdk
Javascript and Typescript SDK for Speechmatics
TypeScript · HTML★ 57↓ 244.1K/月2026年7月26日
MIT2026年7月26日 · 指标 2.10.0
crates.io
73良好健康指数
altunenes/parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Rust★ 378↓ 12.5K/月2026年7月29日
MIT2026年7月29日 · 指标 2.10.0
PyPI · crates.io
67良好健康指数
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
C++ · C★ 197↓ 7,559/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
65良好健康指数
jdepoix/youtube-transcript-api
This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!
Python★ 8,121↓ 28.7M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
npm · crates.io
62中等健康指数
snomiao/otoji
realtime speech ⇄ text, 音を字に — wire mic → STT → translate → speech as an on-device voice graph that spans your devices over WebRTC. Runs in the browser (transformers.js/ONNX/WebGPU); no API keys, nothing leaves the device by default.
TypeScript · Rust★ 2↓ 11.3K/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
Go
59中等健康指数
VoiceBlender/voiceblender
A programmable voice platform: SIP and WebRTC call control, multi-party mixing, recording, TTS/STT, and pluggable AI agents (ElevenLabs, VAPI, Pipecat, Deepgram) — all driven through a REST API, webhooks, and a WebSocket event stream
Go★ 972026年7月20日
MIT2026年7月20日 · 指标 2.10.0
NuGet
57中等健康指数
umlx5h/LLPlayer
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
C#★ 3,9902026年8月5日
GPL-3.02026年8月5日 · 指标 2.10.0
Maven
51中等健康指数
CrispStrobe/CrisperWeaver
On-device speech-to-text Flutter app powered by CrispASR (ggml / Whisper) — offline, multi-platform, AGPL-3.0.
Dart★ 412026年7月30日
AGPL-3.02026年7月30日 · 指标 2.10.0
npm
41薄弱健康指数
surajmandalcell/asrpro
AI powered desktop transcription app with real time speech recognition, file transcription, global hotkeys, and SRT subtitle export.
TypeScript · JavaScript★ 42026年8月1日
无许可证2026年8月1日 · 指标 2.10.0
PyPI · npm · Go +2
28存在风险健康指数
alphacep/vosk-api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Jupyter Notebook★ 15K↓ 776.1K/月2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
23存在风险健康指数
wiseman/py-webrtcvad
Python interface to the WebRTC Voice Activity Detector
C · C++★ 2,492↓ 451.3K/月2026年7月21日
自定义许可证2026年7月21日 · 指标 2.10.0
PyPI
11危急健康指数
abhirooptalasila/AutoSub
A CLI script to generate subtitle files (SRT/VTT/TXT) for any video using either DeepSpeech or Coqui
Python★ 6512026年8月4日
MIT2026年8月4日 · 指标 2.10.0