Examples for ICASSP2024 paper “StemGen: A music generation model that listens”.
Directory
Search results
Published directory entries matching your search.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Professional AI voice generator for real production workflows, with free testing and flexible integration for studios and media teams.
Works with speech, voice, music or other audio using machine-learning models.
APIs for messaging, voice, and phone verification.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
Works with speech, voice, music or other audio using machine-learning models.
A simple notebook demonstrating prompt-based music generation via Mubert API.
Tts for Lojban using VITS TTS models.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
High-quality text-to-speech and voice recognition.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Personalized soundscapes to help you focus, relax, and sleep. Backed by neuroscience.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Free speech-to-text tool for content creators that accurately transcribes audio & video files up to 2GB.
CMU Sphinx uses AI for speech, voice, transcription, music, or other audio workflows.
Text-to-speech solutions with character.
Balabolka is a text-to-speech application (freeware).
Text-to-Audio Generation with Latent Diffusion Models - Speech Research.
(USA) A API company for advanced Speech-to-Text, offering highly accurate transcription, summarization, and audio intelligence.