PyTorch Implementation of No Token Left Behind: Explainability-Aided Image Classification and Generation.
Directory
Search results
Published directory entries matching your search.
Browser extension for privacy-preserving tokens that can reduce repeated CAPTCHA challenges.
Tokenizers for Natural Language Processing in Julia.
GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools.
Smart code context extractor for AI assistants with accurate token counting and budget management.
A C++ library for unsupervised text tokenization and detokenization, widely used in modern NLP models.
Golang implementation of Punkt sentence tokenizer.
GitHub secret scanner that finds exposed tokens, keys, and sensitive strings in public code.
Pure-Rust tokenizer for GGUF models, compatible with llama.cpp tokenization.
Generate tiktok signature token using node
Minimize LLM token complexity to save API costs and model computations.
Open-source DMS (data management system) for powering data hubs and data portals.
Kuwala is the no-code data platform for BI analysts and engineers enabling you to build powerful analytics workflows. We are set out to bring state-of-the-art data engineering tools you love, such as Airbyte, dbt, or Great Expectations together in one intuitive interface built with React Flow. In addition we provide third-party data into data science models and products with a focus on geospatial data. Currently, the following data connectors are available worldwide: a) High-resolution demograph
A federated, open-source data catalog for all your big data and small data.
Open-source project that applies AI to data analysis, extraction, visualization, or business intelligence.
A repository of useful data science prompts for ChatGPT.
Template repository for data science lifecycle project.
Data Version Control - Git for Data & Models - ML Experiments Management.
A lightweight set of tools for loading and sharing data in data science projects.
Open-source project that applies AI to data analysis, extraction, visualization, or business intelligence.
Pandas API on Apache Spark. Makes data scientists more productive when interacting with big data.
Fast and easy data exploration by automating the visualization and data analysis process.
MeTA: ModErn Text Analysis is a C++ Data Sciences Toolkit that facilitates mining big text data.
CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.