Tab-Distillation: Impacts of Dataset Distillation on Tabular Data For Outlier Detection

2Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Dataset distillation aims to replace large training sets with significantly smaller synthetic sets while preserving essential information. This method reduces the training costs of advanced deep learning models and is widely used in the image domain. Among various distillation methods, "Dataset Condensation with Distribution Matching (DM)"stands out for its low synthesis cost and minimal hyperparameter tuning. Due to its computationally economical nature, DM is applicable to realistic scenarios, such as industries with large tabular datasets. However, its use in tabular data has not been extensively explored. In this study, we apply DM to tabular datasets for outlier detection. Our findings show that distillation effectively addresses class imbalance, a common issue in these datasets. The synthetic datasets offer better sample representation and class separation between inliers and outliers. They also maintain high feature correlation making them resilient against feature pruning. Classification models trained on these distilled datasets perform faster and better that will enhance outlier detection in industries that rely on tabular data.

Cite

CITATION STYLE

APA

Herurkar, D., Raue, F., & Dengel, A. (2024). Tab-Distillation: Impacts of Dataset Distillation on Tabular Data For Outlier Detection. In ICAIF 2024 - 5th ACM International Conference on AI in Finance (pp. 804–812). Association for Computing Machinery, Inc. https://doi.org/10.1145/3677052.3698660

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free