Works with speech, voice, music or other audio using machine-learning models.
Directory
Search results
Published directory entries matching your search.
Personalized soundscapes to help you focus, relax, and sleep. Backed by neuroscience.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Fast inference engine for whisper in C++ using CTranslate2.
結婚式スピーチ・弔辞・年賀状・退職挨拶・お礼状・お詫び文など、冠婚葬祭と日常の手紙・挨拶文をAIが作成します。シチュエーションを選んで項目を埋めるだけで、そのまま使える文面が完成。82種類のツールを登録不要・完全無料で今すぐ使えます。.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
Free production platform for text-to-image generation using Nano Banana V2 model.
Port of OpenAI's Whisper model in C/C++. It can be executed locally.
Good Tape uses AI for speech, voice, transcription, music, or other audio workflows.
We are a community-driven organization releasing open-source generative audio tools to make music production more accessible and fun for everyone.
1. Create Images: Generate images from text prompts using Gemini 2.0 Flash.
Creates or edits images with generative models and visual controls.
Showcase streaming text generation using the @huggingface/inference JS lib.
Creates or edits images with generative models and visual controls.
Img To Music uses AI for speech, voice, transcription, music, or other audio workflows.
Distribution, publishing, funding, marketing, and a hands-on team for independent artists. We.
High-quality text-to-speech and voice recognition.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
A curated list of resources of audio-driven talking face generation.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Works with speech, voice, music or other audio using machine-learning models.
Letterfork uses AI for speech, voice, transcription, music, or other audio workflows.
The state of the art AI image generation engine.
Tts for Lojban using VITS TTS models.