Abstract
This paper presents a statistics-based and language independent unsupervised approach for clustering possible named entities. We describe and motivate the features and statistical filters used by our clustering process. Using the Model-Based Clustering Analysis software we obtained different clusters of named entities. The method was applied to Bulgarian and English. For some clusters, precision is close to 100%; this helps human validation and saves time. Other clusters still need further refinement. Based on the obtained clusters, it is possible to classify new named entities.
Cite
CITATION STYLE
Ferreira Da Silva, J. F., Kozareva, Z., & Lopes, J. G. P. (2004). Cluster analysis and classification of named entities. In Proceedings of the 4th International Conference on Language Resources and Evaluation, LREC 2004 (pp. 321–324). European Language Resources Association (ELRA). https://doi.org/10.63317/3xbgqf5airh3
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.