Tabular Data Classification and Regression: XGBoost or Deep Learning with Retrieval-Augmented Generation

14Citations
Citations of this article
69Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Tabular data is the most prevalent form of structured data, necessitating robust models for classification and regression tasks. Traditional models like eXtreme Gradient Boosting (XGBoost) have gained popularity for their strong performance, while deep learning models such as Tabular Retrieval-Augmented Generation (TabR) and TabNet offer innovative approaches. TabR uniquely employs Retrieval-Augmented Generation (RAG) to reduce uncertainty and enhance predictive accuracy, whereas TabNet relies on a sequential attention mechanism without incorporating RAG. This study systematically compares TabR and TabNet in classification and regression tasks using benchmark datasets, with evaluations based on accuracy, Area Under the Curve (AUC), Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R2 (Coefficient of Determination). The results reveal that TabR, with its RAG component, outperforms XGBoost in classification, effectively managing uncertainty. However, in regression tasks, XGBoost continues to excel over TabR. Meanwhile, TabNet performs comparably but lacks the performance enhancement provided by RAG in TabR. These findings highlight the potential of RAG in deep learning models for tabular data classification and suggest areas for further exploration in improving regression performance.

Cite

CITATION STYLE

APA

Pasaribu, J., Yudistira, N., & Firdaus Mahmudy, W. (2024). Tabular Data Classification and Regression: XGBoost or Deep Learning with Retrieval-Augmented Generation. IEEE Access, 12, 191719–191732. https://doi.org/10.1109/ACCESS.2024.3518205

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free