Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Applio is a GitHub repository or organization with source code, releases, documentation, or project resources.
Ace Step 1.5 is a GitHub repository or organization with source code, releases, documentation, or project resources.
Mmaudio is a GitHub repository or organization with source code, releases, documentation, or project resources.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Port of OpenAI's Whisper model in C/C++. opensource.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
An offline speech recognition toolkit with C++ support, designed for low-resource devices and multiple languages.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
Works with speech, voice, music or other audio using machine-learning models.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Works with speech, voice, music or other audio using machine-learning models.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
A simple notebook demonstrating prompt-based music generation via Mubert API.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Fast inference engine for whisper in C++ using CTranslate2.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.