NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
Directory
Search results
Published directory entries matching your search.
Otter AI Meeting Agent supports real-time transcription, live chat, automated summaries, insights, and action items.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
APIs for messaging, voice, and phone verification.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
Discover amazing ML apps made by the community.
MusicGen uses AI for speech, voice, transcription, music, or other audio workflows.
MuseGen uses AI for speech, voice, transcription, music, or other audio workflows.
An AI-powered voiceover tool that provides realistic voices for videos, podcasts, and presentations.
Murf AI uses AI for speech, voice, transcription, music, or other audio workflows.
Works with speech, voice, music or other audio using machine-learning models.
A simple notebook demonstrating prompt-based music generation via Mubert API.
"transforming the future of music creation".
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
Tts for Lojban using VITS TTS models.
Letterfork uses AI for speech, voice, transcription, music, or other audio workflows.
Works with speech, voice, music or other audio using machine-learning models.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
High-quality text-to-speech and voice recognition.
Distribution, publishing, funding, marketing, and a hands-on team for independent artists. We.
Img To Music uses AI for speech, voice, transcription, music, or other audio workflows.
We are a community-driven organization releasing open-source generative audio tools to make music production more accessible and fun for everyone.
Good Tape uses AI for speech, voice, transcription, music, or other audio workflows.