Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Directory
Search results
Published directory entries matching your search.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
結婚式スピーチ・弔辞・年賀状・退職挨拶・お礼状・お詫び文など、冠婚葬祭と日常の手紙・挨拶文をAIが作成します。シチュエーションを選んで項目を埋めるだけで、そのまま使える文面が完成。82種類のツールを登録不要・完全無料で今すぐ使えます。.
Fast inference engine for whisper in C++ using CTranslate2.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Personalized soundscapes to help you focus, relax, and sleep. Backed by neuroscience.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Discover amazing ML apps made by the community.
Works with speech, voice, music or other audio using machine-learning models.
Free speech-to-text tool for content creators that accurately transcribes audio & video files up to 2GB.
Architecture of voice AI, from speech recognition to emotional intelligence, and learn how to build, scale, and evaluate them.
CMU Sphinx uses AI for speech, voice, transcription, music, or other audio workflows.
Text-to-speech solutions with character.
Works with speech, voice, music or other audio using machine-learning models.
Generate full songs with AI for free. Describe your idea — Boppy writes lyrics and creates a complete track in minutes. No signup, no credit card.
Bespoke Sounds uses AI for speech, voice, transcription, music, or other audio workflows.
Balabolka is a text-to-speech application (freeware).
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Text-to-Audio Generation with Latent Diffusion Models - Speech Research.
AudioCraft is a single-stop code base for all your generative audio needs: music, sound effects, and compression after training on raw audio signals.
(USA) A API company for advanced Speech-to-Text, offering highly accurate transcription, summarization, and audio intelligence.