Spark is a fast and general engine for large-scale data processing.
Directory
Search results
Published directory entries matching your search.
Massively parallel self-organizing maps: accelerate training on multicore CPUs, GPUs, and clusters, has python API.
A JavaScript library aimed at visualizing graphs of thousands of nodes and edges.
A python visualization library based on matplotlib.
Basic sampling algorithms for Julia.
A javascript library containing a collection of least squares fitting methods for finding a trend in a set of data.
Redash applies AI to data analysis, extraction, visualization, or business intelligence.
Julia package for loading many of the data sets available in R.
A python framework to transform natural language questions to queries in a database query language.
Simple plotting for Python. Wrapper for D3xterjs; easily render charts in-browser.
A Techie Tech Blog on Master Data Management And Every Buzz Surrounding It.
Portable annotation tool for creating labeled datasets.
The polyglot notebook with first-class Scala support.
Fast DataFrame library for Rust and Python, designed as a faster alternative to Pandas.
CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.
A Clojure/Clojurescript notebook application/-library based on Gorilla-REPL.
Simple, realtime visualization of neural network training performance.
Clojure API wrapping Python's Pandas library.
Create HTML profiling reports from pandas DataFrame objects.
A library providing high-performance, easy-to-use data structures and data analysis tools.
Open-source project that applies AI to data analysis, extraction, visualization, or business intelligence.
Cleansing, pre-processing, feature engineering, exploratory data analysis and easy ML with PySpark backend.
Bring multiple data streams into one dashboard.
Distributed, masterless, high performance, fault tolerant data processing. Written entirely in Clojure.