Mining revision log of language learning SNS for automated Japanese error correction

2Citations
Citations of this article
40Readers
Mendeley users who have this article in their library.

Abstract

Recently, natural language processing research has begun to pay attention to second language learning. However, it is not easy to acquire a large-scale learners' corpus, which is important for a research for second language learning by natural language processing. We present an attempt to extract a large-scale Japanese learners' corpus from the revision log of a language learning social network service.This corpus is easy to obtain in large-scale, covers a wide variety of topics and styles, and can be a great source of knowledge for both language learners and instructors. We also demonstrate that the extracted learners' corpus of Japanese as a second language can be used as training data for learners' error correction using a statistical machine translation approach.We evaluate different granularities of tokenization to alleviate the problem of word segmentation errors caused by erroneous input from language learners.We propose a character-based SMT approach to alleviate the problem of er oneous input from language learners.Experimental results show that the character-based model outperforms the word-based model when corpus size is small and test data is written by the learners whose L1 is English. © The Japanese Society for Artificial Intelligence 2013.

Cite

CITATION STYLE

APA

Mizumoto, T., Komachi, M., Nagata, M., & Matsumoto, Y. (2013). Mining revision log of language learning SNS for automated Japanese error correction. Transactions of the Japanese Society for Artificial Intelligence, 28(5), 420–432. https://doi.org/10.1527/tjsai.28.420

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free