Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Zyphra uses AI for speech, voice, transcription, music, or other audio workflows.
Turn live audio into stunning visuals.
Creates or edits images with generative models and visual controls.
Helps find, analyze or synthesize information for research and knowledge discovery.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
Curate and annotate vision, audio, and LLM datasets, track experiments, and manage models on a single platform.
A large Dataset of synchronised Audio, LyrIcs and vocal notes.
Host inference APIs, bulk inference and fine tune text, vision, audio and multi-modal models.
Weights & Biases, developer tools for machine learning.
ChatGPT, Claude, Gemini and 350+ AI models in one all-in-one AI workspace. Images, video, audio, code — one subscription, no juggling.
Your Personal AI Library — AI summaries, Video Insights, audiobooks, and podcasts. Read less, know more from any book.
An experimental physical interface for the NSynth machine learning algorithm.
Nuance applies AI to workplace tasks, meetings, sales, marketing, or productivity.
An open-source tool for recording screen and audio activity with AI-powered search, automations, and support for local LLMs. opensource.
Multimodal AI video generation model (text, image, video, audio).
SkillBoss supports AI agents, automated workflows, orchestration, or delegated tasks.
Create professional videos using just a still image with text or audio powered by AI.
Create AI generated videos from text with the most advanced AI avatars and voiceovers in 160+ languages. Try our free AI video generator now!
TwelveLabs delivers enterprise video AI powered by multimodal intelligence. Search, analyze, and understand video across vision, audio, and language.