Abstract
Tabular data is the most prevalent form of structured data, necessitating robust models for classification and regression tasks. Traditional models like eXtreme Gradient Boosting (XGBoost) have gained popularity for their strong performance, while deep learning models such as Tabular Retrieval-Augmented Generation (TabR) and TabNet offer innovative approaches. TabR uniquely employs Retrieval-Augmented Generation (RAG) to reduce uncertainty and enhance predictive accuracy, whereas TabNet relies on a sequential attention mechanism without incorporating RAG. This study systematically compares TabR and TabNet in classification and regression tasks using benchmark datasets, with evaluations based on accuracy, Area Under the Curve (AUC), Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R2 (Coefficient of Determination). The results reveal that TabR, with its RAG component, outperforms XGBoost in classification, effectively managing uncertainty. However, in regression tasks, XGBoost continues to excel over TabR. Meanwhile, TabNet performs comparably but lacks the performance enhancement provided by RAG in TabR. These findings highlight the potential of RAG in deep learning models for tabular data classification and suggest areas for further exploration in improving regression performance.
Author supplied keywords
Cite
CITATION STYLE
Pasaribu, J., Yudistira, N., & Firdaus Mahmudy, W. (2024). Tabular Data Classification and Regression: XGBoost or Deep Learning with Retrieval-Augmented Generation. IEEE Access, 12, 191719–191732. https://doi.org/10.1109/ACCESS.2024.3518205
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.