Abstract
K-means clustering is a fundamental tool in data mining, yet its scalability and efficacy decline when faced with massive datasets. In this work, we introduce BiModalClust, a novel clustering algorithm that leverages a bimodal optimization paradigm to overcome these challenges. Our approach simultaneously optimizes two interdependent modalities: the input data stream and the neighborhood structure of the solution landscape, which emerges from iterative restrictions of the Minimum Sum-of-Squares Clustering (MSSC) objective function to sampled subsets of the data. By integrating the Variable Neighborhood Search (VNS) metaheuristic, we systematically explore and refine these landscapes through dynamic reinitialization of degenerate centroids and adaptive exploration of expanding neighborhoods. This dual-stream optimization not only transforms traditional local search into a more global and robust process but also ensures computational scalability and precision. Extensive experimentation on diverse real-world datasets demonstrates that BiModalClust achieves superior clustering performance among K-means-based methods in big data environments.
Author supplied keywords
Cite
CITATION STYLE
Mussabayev, R., & Mussabayev, R. (2025). BiModalClust: Fused Data and Neighborhood Variation for Advanced K-Means Big Data Clustering. Applied Sciences (Switzerland), 15(3). https://doi.org/10.3390/app15031032
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.