Applio uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Audio generation using diffusion models, in PyTorch.
(USA) A API company for advanced Speech-to-Text, offering highly accurate transcription, summarization, and audio intelligence.
Audacity Effects uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Audio.Z.AI uses AI for speech, voice, transcription, music, or other audio workflows.
AudioArena uses AI for speech, voice, transcription, music, or other audio workflows.
AudioCraft is a single-stop code base for all your generative audio needs: music, sound effects, and compression after training on raw audio signals.
Text-to-Audio Generation with Latent Diffusion Models - Speech Research.
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Balabolka is a text-to-speech application (freeware).
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Bespoke Sounds uses AI for speech, voice, transcription, music, or other audio workflows.
Boomy uses AI for speech, voice, transcription, music, or other audio workflows.
Generate full songs with AI for free. Describe your idea — Boppy writes lyrics and creates a complete track in minutes. No signup, no credit card.
Theano based library for deep and recurrent neural networks.
Works with speech, voice, music or other audio using machine-learning models.
Cartesia uses AI for speech, voice, transcription, music, or other audio workflows.
Text-to-speech solutions with character.
Chatterbox uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
CMU Sphinx uses AI for speech, voice, transcription, music, or other audio workflows.
A comparative framework for multimodal recommender systems with a focus on models leveraging auxiliary data.