Abstract
Recent improvement in natural language understanding research can be attributed to the availability of large scale datasets. Those datasets are mainly in English. In this work, we develop a web crawler with the purpose of extracting Indonesian news content from the DetikNews website and building a large dataset of texts. The web crawler is developed by following the waterfall model using Python, Scrapy, and BeautifulSoup4. It collects more than 790k news from DetikNews, spanning from 2011 to 2020, which consists of a total number of more than 190 million words, almost 2 million unique words, and more than 14 million sentences.
Cite
CITATION STYLE
Hendryli, J., & Mawardi, V. C. (2020). Development of web crawler to build Indonesian text corpus. In IOP Conference Series: Materials Science and Engineering (Vol. 1007). IOP Publishing Ltd. https://doi.org/10.1088/1757-899X/1007/1/012043
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.