Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
USE BAZEL VERSION=5.0.0./bazelisk-linux-amd64 build wavegru mod -c opt --copt=-march=native.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
Port of OpenAI's Whisper model in C/C++. opensource.
WolframTones uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Zyphra uses AI for speech, voice, transcription, music, or other audio workflows.
Creates or edits images with generative models and visual controls.
ChatGPT, Claude, Gemini and 350+ AI models in one all-in-one AI workspace. Images, video, audio, code — one subscription, no juggling.
Your Personal AI Library — AI summaries, Video Insights, audiobooks, and podcasts. Read less, know more from any book.
Multimodal AI video generation model (text, image, video, audio).
Creates, rewrites or summarizes written content with language models.
Create professional videos using just a still image with text or audio powered by AI.
Create AI generated videos from text with the most advanced AI avatars and voiceovers in 160+ languages. Try our free AI video generator now!
Based AI creates or edits images with generative AI and text-based controls.
Turn text, scripts, and blog posts into videos with 2,000+ AI voices in 80+ languages. Free AI video generator - no camera, no editing skills needed.
1. Create Images: Generate images from text prompts using Gemini 2.0 Flash.
Photo editing with AI beautification.
Create stunning AI images and videos instantly. Professional AI generator, upscaler, and editing tools. Transform text into high-quality visuals.
Built-in templates for generating or editing any pictures. Moreover, you can create your own design.
Web and mobile photo and video filters and editing tools.
AI-generated visualization prototyping and editing platform, support 2D, 3D models, combined with LLM(Large Language Model) for quick editing.
Turn live audio into stunning visuals.
Helps find, analyze or synthesize information for research and knowledge discovery.