Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
-
Updated
Jul 30, 2026 - Python
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
Free, open source voice dictation for macOS. On-device transcription with Apple's Speech framework. No cloud, no API keys, no account.
Terminal voice-to-text TUI — Qwen3-ASR-1.7B on the Apple GPU via MLX (mlx-speech). Fully local, no PyTorch, transcribes in ~1s. macOS Apple Silicon.
Talk. Ink. Push-to-talk dictation for macOS, 100% on-device. Pick your model: Qwen3-ASR, NVIDIA Nemotron or Voxtral, all via Apple MLX.
Voice dictation for the browser — free, private, MIT. Dictate into any web page, or use the pop-out to dictate for any app on your machine.
Local voice-to-text for macOS and iOS. Multilingual (EN/ZH/JP) with Traditional Chinese output. Runs Qwen3-ASR on Apple Silicon via MLX. No cloud, no subscription.
OpenAI-compatible speech-to-text server for nvidia/nemotron-3.5-asr-streaming-0.6b (NeMo). Runs on the DGX Spark / GB10.
Live speech-to-text streaming on Apple Silicon — Qwen3-ASR + Silero VAD + MLX
Sono is on-device dictation for macOS. Parakeet v3 + Apple Intelligence, nothing leaves your Mac.
Private, local-first meeting recorder + transcription, diarization, AI notes, voice dictation & read-aloud for Windows — runs on your own GPU.
🎬 AI subtitle generator: convert video to SRT subtitles locally with NVIDIA NeMo Parakeet-TDT speech-to-text. GPU-accelerated, word-level timestamps, VAD, LLM correction — a fast offline Whisper alternative.
Fully-local speech-to-text dictation. Hold a hotkey, talk, and the transcript lands in the field you're already in — an NVIDIA Parakeet streaming server plus native macOS and Windows clients. Your voice never leaves your LAN.
Fast native C inference engine for speech recognition & translation, llama.cpp-style: streaming + offline ASR, word timestamps, int8/int4 quantization, CPU/Metal/CUDA. Today it runs the best open models (Parakeet, Canary, Nemotron) — built to host more engines tomorrow.
Private, offline voice-to-text for Windows — 3.6% WER, 100% on-device (Tauri + Rust + Python + a finetuned LLM)
Execute high-performance speech recognition and translation on CPU, Metal, or CUDA using this native C inference engine.
Local FastAPI transcription studio: AssemblyAI Universal-2 (99 lang), FFmpeg, yt-dlp, Word/PDF/ZIP export
Skill de Claude Code que transcribe audios y videos a Markdown estructurado con timestamps y diarización, usando Google Gemini. Reemplazo gratuito de ElevenLabs Scribe / Whisper para quien ya paga Gemini.
Add a description, image, and links to the whisper-alternative topic page so that developers can more easily learn about it.
To associate your repository with the whisper-alternative topic, visit your repo's landing page and select "manage topics."