A simple notebook demonstrating prompt-based music generation via Mubert API.
Directory
AI Audio & Speech
Explore AI Audio & Speech resources in AI & Machine Learning.
Works with speech, voice, music or other audio using machine-learning models.
Murf AI uses AI for speech, voice, transcription, music, or other audio workflows.
An AI-powered voiceover tool that provides realistic voices for videos, podcasts, and presentations.
MuseGen uses AI for speech, voice, transcription, music, or other audio workflows.
MusicGen uses AI for speech, voice, transcription, music, or other audio workflows.
Discover amazing ML apps made by the community.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
APIs for messaging, voice, and phone verification.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Otter AI Meeting Agent supports real-time transcription, live chat, automated summaries, insights, and action items.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
Works with speech, voice, music or other audio using machine-learning models.
Read AI uses AI for speech, voice, transcription, music, or other audio workflows.
ReadSpeaker uses AI for speech, voice, transcription, music, or other audio workflows.
Works with speech, voice, music or other audio using machine-learning models.
Replica Studios uses AI for speech, voice, transcription, music, or other audio workflows.
Resemble AI uses AI for speech, voice, transcription, music, or other audio workflows.
Professional AI voice generator for real production workflows, with free testing and flexible integration for studios and media teams.
Rev AI, part of the Rev family, is a developer-first API that delivers industry- accuracy and fast performance at global scale. Click to learn more.
Works with speech, voice, music or other audio using machine-learning models.
Robust Speech Recognition Leaderboard 2022 uses AI for speech, voice, transcription, music, or other audio workflows.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Turn any idea into scroll-stopping Shorts, Reels, and TikToks with AI visuals, studio voiceovers and synced captions.