Abstract
The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a ``nice"" Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in k-means by iteratively applying the Anderson-Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. Experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.
Author supplied keywords
Cite
CITATION STYLE
Alvarado, E., Belton, R., Fischer, E., Lee, K. J., Palande, S., Percival, S., & Purvine, E. (2025). G-Mapper: Learning a Cover in the Mapper Construction. SIAM Journal on Mathematics of Data Science, 7(2), 572–596. https://doi.org/10.1137/24M1641312
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.