Abstract
Efficiently managing lengthy textual data, particularly in online news, is crucial for enhancing the performance of long text classification. This study explores innovative approaches to streamline the Gross Domestic Product (GDP) computation process by utilizing modern data analytics, Natural Language Processing (NLP), and online news sources. Leveraging online news data introduces real-time information, which promises to improve the accuracy and timeliness of economic indicators like GDP. However, handling the complexity of extensive textual data poses a challenge, demanding advanced NLP techniques. This research shifts from traditional word-weight-based methods to keyword-based extractive summarization techniques. These tailored approaches ensure that selected sentences align precisely with specific keywords relevant to the research case, such as GDP growth rate detection. The study emphasizes the necessity of adapting summarization methods to capture information in unique research contexts effectively. Classification results show that the implementation of sentence selection significantly improves classification accuracy. Specifically, there was an average accuracy increase of 0.0226 for machine learning and 0.0164 for transfer learning models. Additionally, in terms of computational efficiency, sentence selection also accelerates processing time during hyperparameter tuning and fine-tuning, as observed using the same computational resources.
Author supplied keywords
Cite
CITATION STYLE
Sholawatunnisa, D. P., & Suadaa, L. H. (2024). OPTIMIZING LONG TEXT CLASSIFICATION PERFORMANCE THROUGH KEYWORD-BASED SENTENCE SELECTION: A CASE STUDY ON ONLINE NEWS CLASSIFICATION FOR INDONESIAN GDP GROWTH-RATE DETECTION. Barekeng, 18(2), 1081–1094. https://doi.org/10.30598/barekengvol18iss2pp1081-1094
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.