Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Deemix Revival is a GitHub repository or organization with source code, releases, documentation, or project resources.
Deezer Mod is a GitHub repository or organization with source code, releases, documentation, or project resources.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Fast inference engine for whisper in C++ using CTranslate2.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
headphones helps download, manage, or archive music, audio, podcasts, or karaoke media.
A curated list of resources of audio-driven talking face generation.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
Lyrics is a GitHub repository or organization with source code, releases, documentation, or project resources.
Mixxx is a GitHub repository or organization with source code, releases, documentation, or project resources.
A simple notebook demonstrating prompt-based music generation via Mubert API.
musescore-downloader helps download, manage, or archive music, audio, podcasts, or karaoke media.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
EPUB to audiobook converter, optimized for Audiobookshelf.