Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
WiFi-Direct-File-Transfer-App helps transfer files directly between devices using local, peer-to-peer, or browser-based sharing.
Cross-platform Twitch Chat application with 3rd-party addon support!
A non-exhaustive collection of third-party clients and mods for Discord.
Kuwala is the no-code data platform for BI analysts and engineers enabling you to build powerful analytics workflows. We are set out to bring state-of-the-art data engineering tools you love, such as Airbyte, dbt, or Great Expectations together in one intuitive interface built with React Flow. In addition we provide third-party data into data science models and products with a focus on geospatial data. Currently, the following data connectors are available worldwide: a) High-resolution demograph
Linux wrapper tool for use with the Steam client for custom launch options and 3rd party programs
TikTok LIVE API for Python: The definitive 3rd-party library to receive livestream events (comments, gifts, etc.) in realtime from TikTok LIVE.
A comprehensive testing and evaluation framework for voice agents across language models, prompts, and agent personas.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
file-transfer is a file transfer or temporary hosting service for uploading files and sharing download links.
lan-file-transfer is a file transfer or temporary hosting service for uploading files and sharing download links.
Transfer is a file transfer or temporary hosting service for uploading files and sharing download links.
Audio generation using diffusion models, in PyTorch.
Fast inference engine for whisper in C++ using CTranslate2.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.