A method for in-depth comparative evaluation: How (dis)similar are outputs of POS taggers, dependency parsers & coreference resolvers really?

1Citations
Citations of this article
77Readers
Mendeley users who have this article in their library.

Abstract

This paper proposes a generic method for the comparative evaluation of system outputs. The approach is able to quantify the pairwise differences between two outputs and to unravel in detail what the differences consist of. We apply our approach to three tasks in Computational Linguistics, i.e. POS tagging, dependency parsing, and coreference resolution. We find that system outputs are more distinct than the (often) small differences in evaluation scores seem to suggest.

Cite

CITATION STYLE

APA

Tuggener, D. (2017). A method for in-depth comparative evaluation: How (dis)similar are outputs of POS taggers, dependency parsers & coreference resolvers really? In 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017 - Proceedings of Conference (Vol. 1, pp. 188–198). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/e17-1018

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free