headphones helps download, manage, or archive music, audio, podcasts, or karaoke media.
Directory
Search results
Published directory entries matching your search.
A curated list of resources of audio-driven talking face generation.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
A simple notebook demonstrating prompt-based music generation via Mubert API.
musescore-downloader helps download, manage, or archive music, audio, podcasts, or karaoke media.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
EPUB to audiobook converter, optimized for Audiobookshelf.
podgrab helps download, manage, or archive music, audio, podcasts, or karaoke media.
qobuz-dl helps download, manage, or archive music, audio, podcasts, or karaoke media.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
scdl helps download, manage, or archive music, audio, podcasts, or karaoke media.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
streamrip helps download, manage, or archive music, audio, podcasts, or karaoke media.
Tidal-Media-Downloader helps download, manage, or archive music, audio, podcasts, or karaoke media.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.