Works with speech, voice, music or other audio using machine-learning models.
Directory
Search results
Published directory entries matching your search.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Works with speech, voice, music or other audio using machine-learning models.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
A comprehensive testing and evaluation framework for voice agents across language models, prompts, and agent personas.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Implementation of Random Forest in Common Lisp.
The OSINT project, the main idea of which is to collect all the possible Google dorks search combinations and to find the information about the specific web-site: common admin panels, the widespread file types and path traversal. The 100% automated.
Open-source agent skill and npm CLI for daily-updated arXiv, PubMed/PMC, and US policy retrieval with deterministic freshness cutoffs. Apache-2.0.
Open-source project that supports software development with code generation, analysis, debugging, or documentation.
Common solutions and tools developed by Google Cloud's Professional Services team. This repository and its contents are not an officially supported Google product.
An alternative Discord client with voice support made with C++ and GTK 3
YumCut - free AI video generator to turn a prompt into ready vertical videos for TikTok, Reels and YouTube Shorts. Auto script, scenes, voiceover, subtitles and watermark. Built with Next.js. Local-first pipeline + templates, batch rendering and API hooks for creators and indie makers. Self-hosted, FFmpeg-ready, multi-language output. Low cost fast
Audio generation using diffusion models, in PyTorch.
Fast inference engine for whisper in C++ using CTranslate2.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Voice-first AI Assistant for online meetings that can actively participate and solve tasks live during the meeting.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
A simple notebook demonstrating prompt-based music generation via Mubert API.