Architecture of voice AI, from speech recognition to emotional intelligence, and learn how to build, scale, and evaluate them.
Directory
AI Audio & Speech
Explore AI Audio & Speech resources in AI & Machine Learning.
Free speech-to-text tool for content creators that accurately transcribes audio & video files up to 2GB.
Works with speech, voice, music or other audio using machine-learning models.
Discover amazing ML apps made by the community.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Personalized soundscapes to help you focus, relax, and sleep. Backed by neuroscience.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Fast inference engine for whisper in C++ using CTranslate2.
結婚式スピーチ・弔辞・年賀状・退職挨拶・お礼状・お詫び文など、冠婚葬祭と日常の手紙・挨拶文をAIが作成します。シチュエーションを選んで項目を埋めるだけで、そのまま使える文面が完成。82種類のツールを登録不要・完全無料で今すぐ使えます。.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Good Tape uses AI for speech, voice, transcription, music, or other audio workflows.
We are a community-driven organization releasing open-source generative audio tools to make music production more accessible and fun for everyone.
Img To Music uses AI for speech, voice, transcription, music, or other audio workflows.
Distribution, publishing, funding, marketing, and a hands-on team for independent artists. We.
High-quality text-to-speech and voice recognition.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Works with speech, voice, music or other audio using machine-learning models.
Letterfork uses AI for speech, voice, transcription, music, or other audio workflows.
Tts for Lojban using VITS TTS models.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
"transforming the future of music creation".