The application of nltk library for python natural language processing in corpus research

50Citations
Citations of this article
175Readers
Mendeley users who have this article in their library.

Abstract

Corpora play an important role in linguistics research and foreign language teaching. At present, the relevant research on the corpus in China mainly uses WordSmith, Antconc and other retrieval tools. NLTK library, which is based on Python language, can provide more flexible and rich research methods, and it can use unified data standards to avoid the trouble of various data type conversion. At the same time, with the help of Python’s numerous third-party libraries, it can make up for the shortcomings of other tools in syntax analysis, graphic rendering, regular expression retrieval and other aspects. In terms of the main links in corpus research, such as text cleaning, word form restoration, part of speech tagging and text retrieval statistics, this paper takes the US presidential inaugural speech in the corpus as an example to show how to use this tool to process the language data, and introduces the application of Python NLTK library in corpus research.

Cite

CITATION STYLE

APA

Wang, M., & Hu, F. (2021). The application of nltk library for python natural language processing in corpus research. Theory and Practice in Language Studies, 11(9), 1041–1049. https://doi.org/10.17507/tpls.1109.09

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free