Text-to-Audio Generation with Latent Diffusion Models - Speech Research.
Directory
Search results
Published directory entries matching your search.
Balabolka is a text-to-speech application (freeware).
Generate full songs with AI for free. Describe your idea — Boppy writes lyrics and creates a complete track in minutes. No signup, no credit card.
Text-to-speech solutions with character.
Free speech-to-text tool for content creators that accurately transcribes audio & video files up to 2GB.
Discover amazing ML apps made by the community.
Personalized soundscapes to help you focus, relax, and sleep. Backed by neuroscience.
Fast inference engine for whisper in C++ using CTranslate2.
Turn text, scripts, and blog posts into videos with 2,000+ AI voices in 80+ languages. Free AI video generator - no camera, no editing skills needed.
結婚式スピーチ・弔辞・年賀状・退職挨拶・お礼状・お詫び文など、冠婚葬祭と日常の手紙・挨拶文をAIが作成します。シチュエーションを選んで項目を埋めるだけで、そのまま使える文面が完成。82種類のツールを登録不要・完全無料で今すぐ使えます。.
Distribution, publishing, funding, marketing, and a hands-on team for independent artists. We.
Voice-first AI Assistant for online meetings that can actively participate and solve tasks live during the meeting.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Tts for Lojban using VITS TTS models.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
Fine-tuned on AMD MI300X for the AMD Developer Hackathon 2026 (Fine-Tuning Track).
"transforming the future of music creation".
A simple notebook demonstrating prompt-based music generation via Mubert API.
An AI-powered voiceover tool that provides realistic voices for videos, podcasts, and presentations.
Discover amazing ML apps made by the community.
Voice AI that turns calls into outcomes - and gets sharper with each one. Our own model, built for the phone. Sub-400ms response. Live in days.
Otter AI Meeting Agent supports real-time transcription, live chat, automated summaries, insights, and action items.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
The AI content team for solo founders. Ships articles in your voice that rank on Google and get cited by ChatGPT, Perplexity, and Gemini. From $29/mo.