Repeatable, atomic and versioned data lake on top of object storage.
Directory
Search results
Published directory entries matching your search.
Modern columnar data format for ML implemented in Rust.
Pretrain computer vision models on unlabeled data for industrial applications.
Collect, aggregate, and visualize a data ecosystem's metadata.
A tensor-based framework for large-scale data computation which is often regarded as a parallel and distributed version of NumPy.
Analyzes structured data and helps produce queries, insights or visualizations.
Algorithm capable of fully capturing the impact of data drift on performance.
Distributed, masterless, high performance, fault tolerant data processing. Written entirely in Clojure.
Bring multiple data streams into one dashboard.
Cleansing, pre-processing, feature engineering, exploratory data analysis and easy ML with PySpark backend.
Pachyderm is a version control system for data.
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
Analyzes structured data and helps produce queries, insights or visualizations.
Materials and IPython notebooks for "Python for Data Analysis" by Wes McKinney, published by O'Reilly Media.
A self-organizing data hub with S3 support.
Julia package for loading many of the data sets available in R.
A javascript library containing a collection of least squares fitting methods for finding a trend in a set of data.
Analyzes structured data and helps produce queries, insights or visualizations.
High performance distributed data processing in NodeJS.
A system for quickly generating training data with weak supervision.
Self Organizing Map written in Python (Uses neural networks for data analysis).
Spark is a fast and general engine for large-scale data processing.
Analyzes structured data and helps produce queries, insights or visualizations.
A data exploration platform designed to be visual, intuitive, and interactive.