Directory

Search results

Published directory entries matching your search.

Search results

53 listings
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 308
github.com

Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.

AI Audio & Speech 157

Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch.

AI Audio & Speech 276
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 208
github.com

Named-entity recognition using neural networks providing state-of-the-art-results.

AI Image Generation 159
github.com

On-device voice dictation for macOS — transcribes a 5-minute clip in 2.8 s; noise-robust, paste at cursor. 99 languages, ~8 MB, no telemetry. MIT.

AI Audio & Speech 161

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 225

A complete object-oriented environment for machine learning in Matlab.

Models & Machine Learning 308
github.com

Works with speech, voice, music or other audio using machine-learning models.

AI Audio & Speech 350
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 316
github.com

An Optimized Speech-to-Text Pipeline for the Whisper Model.

AI Audio & Speech 213
github.com

CLI and GUI for OSINT. Are you very exhibited on the Internet? Check it! Twitter, Tinder, Facebook, Google, Yandex, BOE. It uses facial recognition to provide more accurate results.

Google Communities & Tools 248
github.com

🔭TeleSpot OSINT lookup from Telephone number using DDGR + BING + GOOGLE + DEHASHED and uses pattern recognition correlations. NOW with TelespotX!

Google Communities & Tools 338
github.com

TIKTOD V3 is a bot application designed to automate interactions on Zefoy website, such as increasing views, hearts, followers, and shares on a specified video. The bot uses technologies like Selenium for web automation and OCR (Optical Character Recognition) for solving captchas.

TikTok 335
github.com

Works with speech, voice, music or other audio using machine-learning models.

AI Audio & Speech 192
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 242
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 247
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 167
github.com

A computer vision tool that protects children's video identities during online video conferencing with anonymizing snapchat-like filters and face recognition tracking. It's a different kind of mask, a fun one! :turtle:

Snapchat 242

Accelerates transcription with the combination of OpenAI's Whisper Large v2, HF Transformers, Optimum, and flash attention.

AI Audio & Speech 184
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 160
github.com

Open-source project that uses AI for speech, voice, transcription, music, or other audio workflows.

AI Audio & Speech 266