Transparent Data Preprocessing for Machine Learning

9Citations
Citations of this article
31Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Data preprocessing is an important task in machine learning which can significantly improve model outcomes. However, evaluating the impact of data preprocessing is often difficult. There is a need for tools which make it transparent to the user on how certain transformations conducted in preprocessing affect the data. Thus, we propose a vision of a transparency system for data preprocessing that provides insights into data preparation pipelines. Our envisioned system consists of a Python library which enables users to log transformations and processed data. Subsequently, the system generates summaries of the data which was processed in the pipeline and so-called change profiles which capture the changes conducted in each processing step. These abstractions offer insight into the transformations and their effects on data. Additionally, the system includes an user interface where users can interactively discover the implemented pipeline and the changes made during preprocessing. This paper presents an initial concept of such a system. It also examines further challenges related to making preprocessing transparent and discusses potential solutions to address these challenges.

Cite

CITATION STYLE

APA

Strasser, S., & Klettke, M. (2024). Transparent Data Preprocessing for Machine Learning. In HILDA 2024 - Workshop on Human-In-the-Loop Data Analytics Co-located with SIGMOD 2024. Association for Computing Machinery, Inc. https://doi.org/10.1145/3665939.3665960

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free