Works with speech, voice, music or other audio using machine-learning models.
Directory
Search results
Published directory entries matching your search.
Works with speech, voice, music or other audio using machine-learning models.
You can also find more comprehensive list on and There's an AI AI Voice Cloning list.
Fast inference engine for whisper in C++ using CTranslate2.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
A simple notebook demonstrating prompt-based music generation via Mubert API.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
An offline speech recognition toolkit with C++ support, designed for low-resource devices and multiple languages.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
Port of OpenAI's Whisper model in C/C++. opensource.
Audio Auditor is a GitHub repository or organization with source code, releases, documentation, or project resources.
🔥🔥🔥自定义Android相机(仿抖音 TikTok),其中功能包括视频人脸识别贴纸,美颜,分段录制,视频裁剪,视频帧处理,获取视频关键帧,视频旋转,添加滤镜,添加水印,合成Gif到视频,文字转视频,图片转视频,音视频合成,音频变声处理,SoundTouch,Fmod音频处理。 Android camera(imitation Tik Tok), which includes video editor,audio editor,video face recognition stickers, segment recording,video cropping, video frame processing, get the first video frame, key frame, video rotation, add filter Mirror ,add watermark ,add gif to video, add text to video, picture to video, audio and video synthesis, audio change processing
Audio Share is a code project with source, releases, documentation, or setup notes.
audio-analyzer is a code project with source, releases, documentation, or setup notes.
audiobook-dl is a code project with source, releases, documentation, or setup notes.
AudioBookConverter is a code project with source, releases, documentation, or setup notes.
AudioSource is an open source Linux, macOS, or Unix resource with code, tools, or setup notes.
Jellyfin Audio Player is a code project with source, releases, documentation, or setup notes.