github.com
Smart code context extractor for AI assistants with accurate token counting and budget management.
Directory
Published directory entries matching your search.
Smart code context extractor for AI assistants with accurate token counting and budget management.
An open dataset with 30 trillion tokens for training Large Language Models.
A C++ library for unsupervised text tokenization and detokenization, widely used in modern NLP models.
Golang implementation of Punkt sentence tokenizer.
Pure-Rust tokenizer for GGUF models, compatible with llama.cpp tokenization.
StreamingLLM is a technique that can enable language models like Llama-2 to have conversations that span across millions of tokens.
Versatile Multi-concept Personalization in Token Modulation Space.
Minimize LLM token complexity to save API costs and model computations.