Development of web crawler to build Indonesian text corpus

2Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Recent improvement in natural language understanding research can be attributed to the availability of large scale datasets. Those datasets are mainly in English. In this work, we develop a web crawler with the purpose of extracting Indonesian news content from the DetikNews website and building a large dataset of texts. The web crawler is developed by following the waterfall model using Python, Scrapy, and BeautifulSoup4. It collects more than 790k news from DetikNews, spanning from 2011 to 2020, which consists of a total number of more than 190 million words, almost 2 million unique words, and more than 14 million sentences.

Cite

CITATION STYLE

APA

Hendryli, J., & Mawardi, V. C. (2020). Development of web crawler to build Indonesian text corpus. In IOP Conference Series: Materials Science and Engineering (Vol. 1007). IOP Publishing Ltd. https://doi.org/10.1088/1757-899X/1007/1/012043

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free