Abstract
Clustering is the search for those partitions that re?ect the structure of an object set? Traditional clustering algorithms search only a small sub?set of all possible clusterings ?the solution space? and consequently? there is no guarantee that the solution found will be optimal? We report here on the application of Genetic Algorithms ?GAs? ? stochastic search algorithms touted as e?ective search methods for large and complex spaces ? to the problem of clustering? GAs whichhavebeen made applicable to the problem of clustering ?by adapting the representation? ?tness function? and developing suitable evolutionary operators? are known as Genetic Clustering Algorithms ?GCAs?? There are two parts to our investigation of GCAs? ?rst we look at clustering into a given number of clusters? The performance of GCAs on three generated data sets? analysed using ???? di?ering combinations of adaptions? establishes their e?cacy? Choice of adaptions and parameter settings is data set dependent? but comparison between results using generated and real data sets indicate that performance is consistent for similar data sets with the same numberofobj ects? clusters? attributes? and a similar distribution of objects? Generally? group?number representations are better suited to the clustering problem? as are dynamic scaling? elite selection and high mutation rates? Independent generalised models ?tted to the correctness and timing results for eachof the generated data sets produced accurate predictions of the performance of GCAs on similar real data sets? While GCAs can be successfully adapted to clustering? and the method produces results as accurate and correct as traditional methods? our ?ndings indicate that? given a criterion based on simple distance metrics? GCAs provide no advantages over traditional methods? Second? weinvestigate the potential of genetic algorithms for the more general clustering problem? where the number of clusters is unknown? Weshow that only simple modi?cations to the adapted GCAs are needed? Wehavedeveloped a merging operator? which with elite selection? is employed to evolve an initial population with a large numberofclusters toward better clusterings? With regards to accuracy and correctness? these GCAs are more successful than optimisation methods suchas simulated annealing? However? such GCAs can become trapped in local minima in the same manner as traditional hierarchical methods? Such trapping is characterised bythe situation where good ?k????clusterings do not result from our merge operator acting on good k? clusterings? A marked improvement in the algorithm is observed with the addition of a local heuristic.
Cite
CITATION STYLE
Hansohm, J. (2002). Two-mode Clustering with Genetic Algorithms (pp. 87–93). https://doi.org/10.1007/978-3-642-55991-4_9
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.