LearnPADS++ Incremental Inference of Ad Hoc Data Formats
An ad hoc data source is any semi-structured, non-standard data source. The format of such data sources is often evolving and frequently lacking documentation. Consequently, off-the-shelf tools for processing such data often do not exist, forcing analysts to develop their own tools, a costly and time-consuming process. In this paper, the authors present an incremental algorithm that automatically infers the format of large-scale data sources. From the resulting format descriptions, they can generate a suite of data processing tools automatically. The system can handle large-scale or streaming data sources whose formats evolve over time.