Identifying Optimal Data Distributions for Enhanced Data Modeling in Machine Learning

4Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

Understanding how data is distributed is crucial for building accurate models in machine learning and data science projects. In this paper, we explore practical methods to help identify the best-fitting distribution for real-world datasets. We cover visual techniques like histograms and Q-Q plots, as well as statistical tests such as Kolmogorov-Smirnov (KS) and Anderson-Darling (AD). We also look at model evaluation using criteria like Akaike information criterion (AIC) and Bayesian information criterion (BIC) to ensure a good fit. To illustrate these methods, we use the California Housing dataset, showing how wrong assumptions about data distribution can lead to poor model performance. By following the guidelines provided in this paper, data scientists can choose the right distribution, leading to more accurate models, better anomaly detection, and smarter decision-making across different fields.

Cite

CITATION STYLE

APA

Jaradat, Y., Masoud, M., Manasrah, A., Alia, M., Suwais, K. M., & Almanasra, S. (2025). Identifying Optimal Data Distributions for Enhanced Data Modeling in Machine Learning. International Journal of Advances in Soft Computing and Its Applications, 17(2), 121–137. https://doi.org/10.15849/IJASCA.250730.07

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free