Lexical association measures and collocation extraction

  • Pecina P
  • 67

    Readers

    Mendeley users who have this article in their library.
  • 69

    Citations

    Citations of this article.

Abstract

We present an extensive empirical evaluation of collocation extraction methods based on lexical association measures and their combination. The exper- iments are performed on three sets of collocation candidates extracted from the Prague Dependency Treebank with manual morphosyntactic annotation and from the Czech National Corpus with automatically assigned lemmas and part-of-speech tags. The collocation candidates were manually labeled as collocational or non- collocational. The evaluation is based on measuring the quality of ranking the candidates according to their chance to form collocations. Performance of the methods is compared by precision-recall curves and mean average precision scores. The work is focused on two-word (bigram) collocations only. We experiment with bigrams extracted from sentence dependency structure as well as from surface word order. Further, we study the effect of corpus size on the performance of the indi- vidual methods and their combination.

Author-supplied keywords

  • Collocations
  • Evaluation
  • Lexical association measures
  • Multiword expressions

Get free article suggestions today

Mendeley saves you time finding and organizing research

Sign up here
Already have an account ?Sign in

Find this document

Authors

  • Pavel Pecina

Cite this document

Choose a citation style from the tabs below

Save time finding and organizing research with Mendeley

Sign up for free