Python library for data-centric AI and machine learning with messy, real-world data and labels.
Directory
Search results
Published directory entries matching your search.
Eurybia monitors data and model drift over time and securizes model deployment with data validation.
Synthetic tabular data generation using GANs, Diffusion Models, and LLMs with adversarial filtering and privacy metrics.
GitHub collection of APT reports, threat actor notes, and security research documents.
) - DVM documentation and kind registry
Interact your data and environment using the local GPT, no data leaks, 100% privately, 100% security.
Exploit-DB paper archive with security research, exploitation papers, and technical writeups.
Python package for geospatial intelligence workflows, location data, and map-based OSINT tasks.
Google Advanced Data Analytics Projects: Automatidata, Waze, Tiktok and Salifort Motors
Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API.
schema-org helps create or test structured data, schema markup, and rich-result eligibility.
Research archive containing dark-web authorship verification datasets and baseline models.
) - RSS/Atom gateway to Nostr. Live at [https://atomstr.data.haus](https://atomstr.data.haus)
Datamining Discord changes from the JS files
Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform.
DVMCP is a bridge implementation that connects Model Context Protocol (MCP) servers to Nostr's Data Vending Machine ecosystem
Provides a central interface to connect your LLM's with external data.
LlamaIndex is a data framework for your LLM applications.
automates gathering website profiling data into a CSV from the "BuiltWith" or "Wappalyzer" API for tech stack information, technographic data, website
🤖 Curated AI OSINT resources — Google dorks, Shodan queries, GitHub dorks, and techniques to discover exposed LLM endpoints, leaked AI API keys, misconfigured vector databases, and unprotected AI agents
Code and documentation to train Stanford's Alpaca models, and generate the data.
Latest Papers and Datasets on Multimodal Large Language Models, and Their Evaluation.
Google Maps Scraper & Lead Generation Tool. Extract 50+ data points including business emails, phone numbers, and social profiles. Includes enrichment features, API access, and no recurring fees
Simple and fast databaseless PHP blogging platform, and Flat-File CMS