A federated, open-source data catalog for all your big data and small data.
Directory
Search results
Published directory entries matching your search.
A repository of useful data science prompts for ChatGPT.
Template repository for data science lifecycle project.
Data Version Control - Git for Data & Models - ML Experiments Management.
A lightweight set of tools for loading and sharing data in data science projects.
Pandas API on Apache Spark. Makes data scientists more productive when interacting with big data.
CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.
Analyzes structured data and helps produce queries, insights or visualizations.
Analyzes structured data and helps produce queries, insights or visualizations.
NumPy and Pandas interface to Big Data.
Lightweight, extensible data validation library for Python.
Read files from Stata, SAS, and SPSS.
A library to compare Pandas, Polars, and Spark data frames. It provides stats and lets users adjust for match accuracy.
Reproducible data setup for reproducible science.
Library for working with tabular data in Julia.
Collect, clean and visualize your data in Python.
Open-source package for validating ML models & data, with various checks and suites.
Drop-in replacement for Jupyter and an AI-native workspace for modern data teams.
Library of SAS Enterprise Miner process flow diagrams to help you learn by example about specific data mining topics.
Functions and data dependencies for loading various word embeddings.
Clojure Data Visualisation library, based on Statistiker and D3.
Datawrapper An open source data visualization platform helping everyone to create simple, correct and embeddable charts. Also at.
Analyzes structured data and helps produce queries, insights or visualizations.
Analyzes structured data and helps produce queries, insights or visualizations.