Feature Extraction Optimization to Improve Naïve Bayes Accuracy in Sentiment Analysis of Bulukumba Tourism Objects

  • Setiawan D
  • Umar N
  • Nur M
N/ACitations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

Abstrak Penelitian ini menggunakan media sosial (Twitter) dalam penerapan analisis sentimen untuk menentukan tingkat kepuasan masyarakat terhadap objek wisata Bulukumba. Data teks yang tidak terstruktur menjadi tantangan utama dalam analisis sentimen. Untuk itu, implementasi algoritma Naïve Bayes merupakan pendekatan efektif dalam mengatasi tantangan ini, karena kemampuannya dalam menangani data teks dengan baik. Penelitian ini bertujuan untuk mengevaluasi kinerja Multinomial Naïve Bayes melalui pengujian kombinasi nilai parameter Minimum Document Frequency (min-df) dan Maximum Document Frequency (max-df) dalam menentukan tingkat akurasi. Tahapan analisis ini mencakup pengumpulan data dari Twitter terkait objek wisata Bulukumba. Prapemrosesan yang dilakukan meliputi pembersihan data, casefolding, normalisasi teks, tokenisasi, penghapusan stopword, dan stemming. Ekstraksi fitur menggunakan Count Vectorizer dan pembobotan TF-IDF. Proses diakhiri dengan 10-Fold Cross-Validation dengan memecah data menjadi data latih dan data uji untuk klasifikasi analisis sentimen, serta evaluasi dengan menggunakan Confusion Matrix. Dalam penelitian ini terdapat 10 skenario pengujian yang memiliki kombinasi min-df dan max-df yang berbeda. Nilai min-df yang digunakan terdiri dari 0.001, 0.002, 0.005, 0.01, 0.02 dan untuk max-df terdiri dari 0.5, dan 0.8. Hasil dari implementasi Multinomial Naïve Bayes pada pengujian tersebut menunjukkan bahwa peningkatan akurasi klasifikasi pada pengaturan parameter min-df dan max-df yang efektif. Akurasi tertinggi sebesar 0.7910 pada pengujian kombinasi nilai parameter min-df 0.001 dan max-df 0.8. Sementara itu, rata-rata akurasi setiap pengujian didapatkan nilai tertinggi sebesar 0.7272 dengan min-df 0.002 serta max-df masing-masing 0.5 dan 0.8. Abstract This research employs social media (Twitter) to apply sentiment analysis ascertain the degree of public satisfaction with the Bulukumba tourist attraction. Unstructured text data is a major challenge in sentiment analysis. For this reason, implementing the Naïve Bayes algorithm is an effective approach for conquering this challenge because of its ability to handle text data well. This study aims to evaluate the performance of multinomial Naïve Bayes by testing a combination of minimum document frequency (min-df) and maximum document frequency (max-df) parameter values in determining the level of accuracy. This analysis stage includes collecting data from Twitter related to the Bulukumba tourist attraction. Preprocessing carried out includes data cleaning, casefolding, text normalization, tokenization, stopword removal, and stemming. Feature extraction using Count Vectorizer and TF-IDF weighting. The process ends with 10-Fold Cross-Validation by separating the data into training data and test data for sentiment analysis classification, as well as evaluation using the Confusion Matrix. In this research, there are 10 test scenarios with various combinations of min-df and max-df. The values of employed min-df consists of 0.001, 0.002, 0.005, 0.01, 0.02 and max-df consists of 0.5 and 0.8. The results of implementing Multinomial Naïve Bayes in this test show that classification accuracy increases with effective min-df and max-df parameter settings. The greatest accuracy was 0.7910 in testing a combination of min-df parameter values of 0.001 and max-df 0.8.

Cite

CITATION STYLE

APA

Setiawan, D., Umar, N., & Nur, M. A. (2024). Feature Extraction Optimization to Improve Naïve Bayes Accuracy in Sentiment Analysis of Bulukumba Tourism Objects. SISTEMASI, 13(5), 2209. https://doi.org/10.32520/stmsi.v13i5.4580

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free