This project was created with DeepSite.
Directory
Search results
Published directory entries matching your search.
This project was created with DeepSite.
Creates, edits or enhances video and animation with generative models.
Recast Studio uses AI to create, edit, transform, or animate video content.
Gemini Omni prompts adapted from official docs and community-verified testing. Free, copy-paste ready, every prompt cites its sources.
Easy way to create royalty free music.
Turn videos, podcasts & streams into viral clips with AI captions, face-tracking & auto-posting to TikTok, Shorts & Reels. Try ClipSpeedAI for $1.
Kokoro-82M supports machine learning models, deployment, inspection, datasets, or AI development workflows.
Applio is a GitHub repository or organization with source code, releases, documentation, or project resources.
Ace Step 1.5 is a GitHub repository or organization with source code, releases, documentation, or project resources.
AI Jukebox supports machine learning models, deployment, inspection, datasets, or AI development workflows.
Mmaudio is a GitHub repository or organization with source code, releases, documentation, or project resources.
ACE-Step 1.5 supports machine learning models, deployment, inspection, datasets, or AI development workflows.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
VoiceSphere AI: Streamline document handling for PDFs, DOCs, PPTs, videos, texts. Fast, precise answers. #VoiceSphereAI #AIChat #DocumentManagement.
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
AI transcription & translation for audio/video. Two services: transcribe or translate in 50+ languages. 98% accuracy, pay-as-you-go.
Examples for ICASSP2024 paper “StemGen: A music generation model that listens”.
This premium domain name is available for purchase!
An Optimized Speech-to-Text Pipeline for the Whisper Model.
Rev AI, part of the Rev family, is a developer-first API that delivers industry- accuracy and fast performance at global scale. Click to learn more.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.