Cartesia uses AI for speech, voice, transcription, music, or other audio workflows.
Directory
Search results
Published directory entries matching your search.
Robust Speech Recognition Leaderboard 2022 uses AI for speech, voice, transcription, music, or other audio workflows.
Tunisian Speech Recognition uses AI for speech, voice, transcription, music, or other audio workflows.
Documentation for the caret package.
Works with speech, voice, music or other audio using machine-learning models.
Works with speech, voice, music or other audio using machine-learning models.
Face recognition library that recognizes and manipulates faces from Python or from the command line.
The Gesture Recognition Toolkit (GRT) is a cross-platform, open-source, C++ machine learning library designed for real-time gesture recognition.
NeMo Parakeet ASR Models attain strong speech recognition accuracy while being efficient for inference. Available in CTC and RNN-Transducer variants.
This package contains the matlab implementation of the algorithms described in the book Pattern Recognition and Machine Learning by C. Bishop.
AI Career Advisor supports AI agents, automated workflows, orchestration, or delegated tasks.
Career assistant that combines intelligent job searching, resume matching, and cover letter generation using an agent-based architecture.
Caricature Maker creates or edits images with generative AI and text-based controls.
Cartography of generative AI creates or edits images with generative AI and text-based controls.
Cartopy is a Python package designed for geospatial data processing in order to produce maps and other geospatial data analyses.
Data Science funny cartoons from CartoonStock directory - the world.
Vision Zero Report Card applies AI to data analysis, extraction, visualization, or business intelligence.
Generate full songs with AI for free. Describe your idea — Boppy writes lyrics and creates a complete track in minutes. No signup, no credit card.
Architecture of voice AI, from speech recognition to emotional intelligence, and learn how to build, scale, and evaluate them.
Speech recognition toolkit with streaming ASR, VAD, punctuation, speaker diarization, and OpenAI-compatible serving for voice AI applications.
High-quality text-to-speech and voice recognition.
An open-source framework by NVIDIA for building speech AI systems, including automatic speech recognition and text-to-speech. opensource.
Space emoji (emoji-only character allowed).
An offline speech recognition toolkit with C++ support, designed for low-resource devices and multiple languages.