An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
Directory
Search results
Published directory entries matching your search.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
An offline speech recognition toolkit with C++ support, designed for low-resource devices and multiple languages.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Works with speech, voice, music or other audio using machine-learning models.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Works with speech, voice, music or other audio using machine-learning models.
Audio generation using diffusion models, in PyTorch.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Fast inference engine for whisper in C++ using CTranslate2.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.