Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
A comprehensive testing and evaluation framework for voice agents across language models, prompts, and agent personas.
Open source voice chat software for low-latency group communication and self-hosted servers.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
Open Source framework for voice and multimodal conversational AI.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
Port of OpenAI's Whisper model in C/C++. opensource.
↗ external - Open-source dictation that types where you talk.
An alternative Discord client with voice support made with C++ and GTK 3
YumCut - free AI video generator to turn a prompt into ready vertical videos for TikTok, Reels and YouTube Shorts. Auto script, scenes, voiceover, subtitles and watermark. Built with Next.js. Local-first pipeline + templates, batch rendering and API hooks for creators and indie makers. Self-hosted, FFmpeg-ready, multi-language output. Low cost fast
Audio generation using diffusion models, in PyTorch.
Fast inference engine for whisper in C++ using CTranslate2.
Voice-first AI Assistant for online meetings that can actively participate and solve tasks live during the meeting.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.