A boundary-based tokenization technique for extractive text summarization

  • Nnaemeka M Oparauwah
  • Juliet N Odii
  • Ikechukwu I Ayogu
  • et al.
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

The need to extract and manage vital information contained in copious volumes of text documents has given birth to several automatic text summarization (ATS) approaches. ATS has found application in academic research, medical health records analysis, content creation and search engine optimization, finance and media. This study presents a boundary-based tokenization method for extractive text summarization. The proposed method performs word tokenization by defining word boundaries in place of specific delimiters. An extractive summarization algorithm was further developed based on the proposed boundary-based tokenization method, as well as word length consideration to control redundancy in summary output. Experimental results showed that the proposed approach enhanced word tokenization by enhancing the selection of appropriate keywords from text document to be used for summarization.

Cite

CITATION STYLE

APA

Nnaemeka M Oparauwah, Juliet N Odii, Ikechukwu I Ayogu, & Vitalis C Iwuchukwu. (2021). A boundary-based tokenization technique for extractive text summarization. World Journal of Advanced Research and Reviews, 11(2), 303–312. https://doi.org/10.30574/wjarr.2021.11.2.0351

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free