Abstract
Data preprocessing is an important task in machine learning which can significantly improve model outcomes. However, evaluating the impact of data preprocessing is often difficult. There is a need for tools which make it transparent to the user on how certain transformations conducted in preprocessing affect the data. Thus, we propose a vision of a transparency system for data preprocessing that provides insights into data preparation pipelines. Our envisioned system consists of a Python library which enables users to log transformations and processed data. Subsequently, the system generates summaries of the data which was processed in the pipeline and so-called change profiles which capture the changes conducted in each processing step. These abstractions offer insight into the transformations and their effects on data. Additionally, the system includes an user interface where users can interactively discover the implemented pipeline and the changes made during preprocessing. This paper presents an initial concept of such a system. It also examines further challenges related to making preprocessing transparent and discusses potential solutions to address these challenges.
Author supplied keywords
Cite
CITATION STYLE
Strasser, S., & Klettke, M. (2024). Transparent Data Preprocessing for Machine Learning. In HILDA 2024 - Workshop on Human-In-the-Loop Data Analytics Co-located with SIGMOD 2024. Association for Computing Machinery, Inc. https://doi.org/10.1145/3665939.3665960
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.