A simple notebook demonstrating prompt-based music generation via Mubert API.
Directory
Search results
Published directory entries matching your search.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
Port of OpenAI's Whisper model in C/C++. opensource.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Browser extension that helps solve audio CAPTCHA challenges using speech recognition.
Speech to Text and KB input captions for OBS, VRChat, Twitch chat and Discord
Foundational Models for State-of-the-Art Speech and Text Translation.
Python library and CLI tool to interface with Google Translate's text-to-speech API
Python bindings for ZPar, a statistical part-of-speech-tagger, constituency parser, and dependency parser for English.
Generate TikTok Text-to-Speech voices in your browser
Python-zpar - Python bindings for, a statistical part-of-speech-tagger, constituency parser, and dependency parser for English.