Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.
C, C++, and Python tools for named entity recognition and relation extraction.
A simple notebook demonstrating prompt-based music generation via Mubert API.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Named-entity recognition using neural networks providing state-of-the-art-results.
On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
A complete object-oriented environment for machine learning in Matlab.
Works with speech, voice, music or other audio using machine-learning models.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
An Optimized Speech-to-Text Pipeline for the Whisper Model.
CLI and GUI for OSINT. Are you very exhibited on the Internet? Check it! Twitter, Tinder, Facebook, Google, Yandex, BOE. It uses facial recognition to provide more accurate results.
ðŸ”TeleSpot OSINT lookup from Telephone number using DDGR + BING + GOOGLE + DEHASHED and uses pattern recognition correlations. NOW with TelespotX!
TIKTOD V3 is a bot application designed to automate interactions on Zefoy website, such as increasing views, hearts, followers, and shares on a specified video. The bot uses technologies like Selenium for web automation and OCR (Optical Character Recognition) for solving captchas.
Works with speech, voice, music or other audio using machine-learning models.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
A computer vision tool that protects children's video identities during online video conferencing with anonymizing snapchat-like filters and face recognition tracking. It's a different kind of mask, a fun one! :turtle:
Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.
Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.