The MADlib Analytics Library or MAD Skills, the SQL
MADlib is a free, open-source library of in-database analytic methods. It provides an evolving suite of SQL-based algorithms for machine learning, data mining and statistics that run at scale within a database engine, with no need for data import/export to other tools. The goal is for MADlib to eventually serve a role for scalable database systems that is similar to the CRAN library for R: a community repository of statistical methods, this time written with scale and parallelism in mind. In this paper, the authors introduce the MADlib project, including the background that led to its beginnings, and the motivation for its open-source nature.