LTL-UDE @ EmpiriST 2015: Tokenization and PoS Tagging of Social Media Text

10Citations
Citations of this article
75Readers
Mendeley users who have this article in their library.

Abstract

We present a detailed description of our submission to the EmpiriST shared task 2015 for tokenization and part-of-speech tagging of German social media text. As relatively little training data is provided, neither tokenization nor PoS tagging can be learned from the data alone. For tokenization, our system uses regular expressions for general cases and word lists for exceptions. For PoS tagging, adding unsupervised knowledge beyond the available training data is the most important factor for reaching acceptable tagging accuracy. A learning curve experiment shows furthermore that more in-domain training data is very likely to further increase accuracy.

Cite

CITATION STYLE

APA

Horsmann, T., & Zesch, T. (2016). LTL-UDE @ EmpiriST 2015: Tokenization and PoS Tagging of Social Media Text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 120–126). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w16-2615

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free