CHID: A large-scale Chinese IDiom dataset for cloze test

51Citations
Citations of this article
148Readers
Mendeley users who have this article in their library.

Abstract

Cloze-style reading comprehension in Chinese is still limited due to the lack of various corpora. In this paper we propose a large-scale Chinese cloze test dataset ChID, which studies the comprehension of idiom, a unique language phenomenon in Chinese. In this corpus, the idioms in a passage are replaced by blank symbols and the correct answer needs to be chosen from well-designed candidate idioms. We carefully study how the design of candidate idioms and the representation of idioms affect the performance of state-of-the-art models. Results show that the machine accuracy is substantially worse than that of human, indicating a large space for further research.

Cite

CITATION STYLE

APA

Zheng, C., Huang, M., & Sun, A. (2020). CHID: A large-scale Chinese IDiom dataset for cloze test. In ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (pp. 778–787). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p19-1075

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free