Peer-to-peer network of data owners and data scientists who can collectively train AI models using PySyft.
Directory
Search results
Published directory entries matching your search.
schema-org helps create or test structured data, schema markup, and rich-result eligibility.
Research archive containing dark-web authorship verification datasets and baseline models.
automates gathering website profiling data into a CSV from the "BuiltWith" or "Wappalyzer" API for tech stack information, technographic data, website
GitHub awesome list collecting OSINT tools, datasets, training links, and research resources.
50 AI agent skills for affiliate marketing. Research trending content, write data-backed posts, generate infographics, build landing pages, deploy — full flywheel with social intelligence. Works with Claude Code, Pi, ChatGPT, Gemini, Cursor, Windsurf, any AI.
🤖 Curated AI OSINT resources — Google dorks, Shodan queries, GitHub dorks, and techniques to discover exposed LLM endpoints, leaked AI API keys, misconfigured vector databases, and unprotected AI agents
A comprehensive set of fairness metrics for datasets and machine learning models.
Easy way to turn any app into searchable data for LLMs.
Code and documentation to train Stanford's Alpaca models, and generate the data.
Lightweight analytics reporting and publishing tool for Digital Analytics Program's Google Analytics 360 data.
Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler.
Platform for Production Data Science.
Scripts to generate a dataset with static frames from the Arcade Learning Environment.
Automated machine learning for image, text, tabular, time-series, and multi-modal data.
Open-source project that provides AI-assisted learning, courses, tutorials, or study support.
) - A fast web client boilerplate written in C# / Blazor, that uses an in-browser SQLite database.
) - An efficient, scalable, and deduplicated local blob storage that maintains metadata in SQLite. Fully compatible with Blossom, it gives developers a reliable database option for building their own Blossom servers.
Latest Papers and Datasets on Multimodal Large Language Models, and Their Evaluation.
Open-source project that provides machine-learning models, research, training resources, or evaluation tools.
bulk-downloader-for-reddit helps export or download comments, posts, media, or discussion data from online platforms.
An easy-to-use feature store. Optimized for time-series data.
Chiasmodon is an OSINT tool designed to assist in the process of gathering information about a target domain. Its primary functionality revolves around searching for domain-related data, including domain emails, domain credentials, CIDRs , ASNs , and subdomains, the tool also allows users to search Google Play application ID.
Basic functions for clustering data: k-means, dp-means, etc.