Robust Speech Recognition Leaderboard 2022 uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Tunisian Speech Recognition uses AI for speech, voice, transcription, music, or other audio workflows.
screenpipe applies AI to workplace tasks, meetings, sales, marketing, or productivity.
An open-source tool for recording screen and audio activity with AI-powered search, automations, and support for local LLMs. opensource.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Convolutional-Recursive Deep Learning for 3D Object Classification[DEEP LEARNING].
Architecture of voice AI, from speech recognition to emotional intelligence, and learn how to build, scale, and evaluate them.
The industry- digital worker platform for recruitment, sales and research agents. Meet our first worker, Lucy—your AI recruitment partner.
Face recognition library that recognizes and manipulates faces from Python or from the command line.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
The Gesture Recognition Toolkit (GRT) is a cross-platform, open-source, C++ machine learning library designed for real-time gesture recognition.
High-quality text-to-speech and voice recognition.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
This package contains the matlab implementation of the algorithms described in the book Pattern Recognition and Machine Learning by C. Bishop.
Recall uses AI to draft, edit, summarize, or adapt written content.
Recast Studio uses AI to create, edit, transform, or animate video content.
Recraft creates or edits images with generative AI and text-based controls.
The Four Wars of the AI Stack (Dec 2023 Recap) applies AI to data analysis, extraction, visualization, or business intelligence.
Space emoji (emoji-only character allowed).
An offline speech recognition toolkit with C++ support, designed for low-resource devices and multiple languages.
A simple and efficient end-to-end Automatic Speech Recognition (ASR) system from Facebook AI Research.
Beautiful Soup: a library designed for screen-scraping HTML and XML.