Abstract
This paper reports on the generation of unambiguous clusters of from clickthrough data from the MSN search query log (the RFP 2006 dataset). Selections (clickthroughs) by a user from a single query can be assumed to have some semantic relevance, and the URLs coselected in this way be aggregated to form single-sense clusters. When the graphs a single term separate into distinct clusters, the semantics of distinct clusters can be interpreted as disambiguated of URLs. This principle had been tested on smaller more constrained datasets previously, and this paper reports findings from applying a method based on the principle to the 2006 dataset. paper evaluates the proposed coselection method for single-sense clusters against two other methods, with parameters. The evaluation is done both with a human to determine the quality of the clusters generated by the methods, and by a simple "edit distance" analysis to the content difference of the methods. main questions addressed are i) whether it is feasie to single-sense / sense-coherent clusters, and ii) whether, in closed world, it would be feasible to discover ambiguous terms. experimentation showed that sense-coherent clusters were and further indicated that ambiguous terms could be detected from observing small overlap between large clusters. Copyright 2009.
Author supplied keywords
Cite
CITATION STYLE
Smith, G., Brailsford, T., Donner, C., Hooijmaijers, D., Truran, M., Goulding, J., & Ashman, H. (2009). Generating unambiguous URL clusters from web search. In Proceedings of Workshop on Web Search Click Data, WSCD’09 (pp. 28–34). Association for Computing Machinery. https://doi.org/10.1145/1507509.1507514
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.