AI Jukebox supports machine learning models, deployment, inspection, datasets, or AI development workflows.
Directory
Search results
Published directory entries matching your search.
Mmaudio is a GitHub repository or organization with source code, releases, documentation, or project resources.
ACE-Step 1.5 supports machine learning models, deployment, inspection, datasets, or AI development workflows.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
VoiceSphere AI: Streamline document handling for PDFs, DOCs, PPTs, videos, texts. Fast, precise answers. #VoiceSphereAI #AIChat #DocumentManagement.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
AI transcription & translation for audio/video. Two services: transcribe or translate in 50+ languages. 98% accuracy, pay-as-you-go.
Examples for ICASSP2024 paper “StemGen: A music generation model that listens”.
This premium domain name is available for purchase!
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Rev AI, part of the Rev family, is a developer-first API that delivers industry- accuracy and fast performance at global scale. Click to learn more.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
An AI-powered voiceover tool that provides realistic voices for videos, podcasts, and presentations.
A simple notebook demonstrating prompt-based music generation via Mubert API.
We are a community-driven organization releasing open-source generative audio tools to make music production more accessible and fun for everyone.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Fast inference engine for whisper in C++ using CTranslate2.
Free speech-to-text tool for content creators that accurately transcribes audio & video files up to 2GB.
AudioCraft is a single-stop code base for all your generative audio needs: music, sound effects, and compression after training on raw audio signals.
(USA) A API company for advanced Speech-to-Text, offering highly accurate transcription, summarization, and audio intelligence.
AI Wedding Toast uses AI for speech, voice, transcription, music, or other audio workflows.
AI Mastering is an automated online audio mastering service using AI. Free mastering is available.