Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Directory
Search results
Published directory entries matching your search.
Kyutai TTS uses AI for speech, voice, transcription, music, or other audio workflows.
LazyPy uses AI for speech, voice, transcription, music, or other audio workflows.
Tts for Lojban using VITS TTS models.
LOVO uses AI for speech, voice, transcription, music, or other audio workflows.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
"transforming the future of music creation".
Mazmazika uses AI for speech, voice, transcription, music, or other audio workflows.
MMAudio uses AI for speech, voice, transcription, music, or other audio workflows.
Moe TTS uses AI for speech, voice, transcription, music, or other audio workflows.
Mubert uses AI for speech, voice, transcription, music, or other audio workflows.
A simple notebook demonstrating prompt-based music generation via Mubert API.
An AI-powered voiceover tool that provides realistic voices for videos, podcasts, and presentations.
MusicFX uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
MusicGPT uses AI for speech, voice, transcription, music, or other audio workflows.
Discover amazing ML apps made by the community.
MVSEP uses AI for speech, voice, transcription, music, or other audio workflows.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
APIs for messaging, voice, and phone verification.
Ondoku uses AI for speech, voice, transcription, music, or other audio workflows.
OpenAI.fm uses AI for speech, voice, transcription, music, or other audio workflows.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.