Abstract
The marketing strategy is very important to follow the culture of visitors or buyers because it is closely related to people's income levels. A number of visitor data are a data mining model that can extract information to determine the characteristics of each data. The purpose of this research is to compare distance measurements using the k-means clustering algorithm to see the optimal k value and the required time complexity. Using the K-Means clustering method with Euclidean, Manhattan, Minkowsky, Chebyshev, and Canberra distances to calculate the characteristic values of each object. Determining the value of k using the Elbow model which is formed from the Sum of Square Error (SSE) also considers the Mean of Square Error (MSE) value. The results showed that the Euclidean, Manhattan, Minkowsky, and Chebyshev distances can provide the right grouping so that they become an alternative to the Euclidean distance where the time needed by the Manhattan distance is 1.70 seconds faster than the Euclidean distance of 1.78 seconds, Minkowsky distance 1.82 seconds, Chebyshev distance 2.30 seconds and Canberra distance of 2.48 seconds. In conclusion, Euclidean, Manhattan, Minkowsky and Chebyshev distances can be used to measure closeness values between objects with good accuracy while Canberra distance cannot provide precise accuracy. The research resulted in five groups with different characteristics of income and expenses so that they can be used as a standard for developing marketing strategies.
Cite
CITATION STYLE
Sakur, S. B. H. (2023). Analisis Perbandingan Pengukuran Jarak pada Algoritme K-Means Berbasis Sum of Square Error. Progresif: Jurnal Ilmiah Komputer, 19(2), 505. https://doi.org/10.35889/progresif.v19i2.1276
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.