asr — independent software & tools
-
aurum
Speech both ways. On-device by default. Local STT + TTS CLI, aurum-core, aurum-ffi.
-
transcria
Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies, docker
-
docker-talkies
OpenAI-compatible audio server in Docker. 7 ASR backends (Whisper, Distil-Whisper, Parakeet, Canary, Canary-Qwen) + 2 TTS engines (Kokoro, Qwen3-TTS voice cloning). Single /v1/audio/{transcriptions,speech,voices} surface. CPU + CUDA images. Hot model swap. MCP server built in.
-
franken_whisper
Agent-first Rust ASR orchestration stack: Bayesian backend routing across whisper.cpp/insanely-fast-whisper/whisper-diarization, real-time NDJSON streaming, SQLite persistence, TTY audio transport, conformance harness. 107K lines, 2000+ tests, zero unsafe code.
-
JianYan
🎤 Transform speech to text on Windows with fast, local AI processing. Enjoy seamless recording and automatic integration for effective communication.
-
vidknot
VidkNot Video Knowledge, Knotted. Convert video links to structured notes, supports Feishu, Yuque, Notion, Obsidian. | 一键将视频链接转换为结构化笔记, 支持飞书、语雀、Notion、Obsidian 等多平台存储.
-
captainslog-whisper
Convert your voice to text locally using Whisper without sending data to the cloud or requiring accounts.