Abstract
We present a large-scale Native Language Identification (NLI) experiment on new data, with a focus on cross-corpus evaluation to identify corpus- and genre-independent language transfer features. We test a new corpus and show it is comparable to other NLI corpora and suitable for this task. Cross-corpus evaluation on two large corpora achieves good accuracy and evidences the existence of reliable language transfer features, but lower performance also suggests that NLI models are not completely portable across corpora. Finally, we present a brief case study of features distinguishing Japanese learners' English writing, demonstrating the presence of cross-corpus and cross-genre language transfer features that are highly applicable to SLA and ESL research.
Cite
CITATION STYLE
Malmasi, S., & Dras, M. (2015). Large-scale Native Language Identification with cross-corpus evaluation. In NAACL HLT 2015 - 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 1403–1409). Association for Computational Linguistics (ACL). https://doi.org/10.3115/v1/n15-1160
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.