Sa`adah, Layla Mazidatus (2026) Analisis Sentimen Berbasis Aspek pada Ulasan Aplikasi Sapawarga Jawa Barat Super Apps Menggunakan LDA dan IndoBERT. Undergraduate thesis, UPN Veteran Jawa Timur.
|
Text (Cover)
22082010105-cover.pdf Download (1MB) |
|
|
Text (Bab 1)
22082010105-bab1.pdf Download (164kB) |
|
|
Text (Bab 2)
22082010105-bab2.pdf Restricted to Repository staff only until 20 July 2029. Download (626kB) |
|
|
Text (Bab 3)
22082010105-bab3.pdf Restricted to Repository staff only until 20 July 2029. Download (636kB) |
|
|
Text (Bab 4)
22082010105-bab4.pdf Restricted to Repository staff only until 20 July 2029. Download (2MB) |
|
|
Text (Bab 5)
22082010105-bab5.pdf Download (138kB) |
|
|
Text (Daftar Pustaka)
22082010105-daftarpustaka.pdf Download (158kB) |
|
|
Text (Lampiran)
22082010105-lampiran.pdf Restricted to Repository staff only until 20 July 2029. Download (805kB) |
Abstract
User reviews of the Sapawarga Jawa Barat SuperApps provide valuable insights into users' experiences and opinions regarding the various features and services offered by the application. However, conventional sentiment analysis approaches are unable to associate sentiment with the specific aspects discussed in each review. This study applies Aspect-Based Sentiment Analysis (ABSA) by utilizing Latent Dirichlet Allocation (LDA) to identify the main aspects and compares the performance of Naïve Bayes, Support Vector Machine (SVM), and IndoBERT for aspect-based sentiment classification. The dataset was collected through web scraping from the Google Play Store and Apple App Store, resulting in 6,012 reviews, which were filtered to obtain 4,386 valid reviews. The research process consisted of text preprocessing, topic modeling using LDA, data annotation, class imbalance handling, and model development using Naïve Bayes, SVM, and IndoBERT for comparative evaluation. The LDA topic modeling identified three main aspects, namely login and account access issues, application performance, and vehicle tax payment features. The optimal number of topics was determined to be three based on the highest coherence score of 0.4247. Among the machine learning models, Naïve Bayes consistently outperformed SVM, while the Random Oversampling (ROS) and Synthetic Minority Oversampling Technique (SMOTE) methods achieved relatively comparable performance. The best machine learning model was Naïve Bayes with TF-IDF unigram and bigram features combined with ROS, achieving an Accuracy of 87.88%, Precision of 79.15%, Recall of 81.38%, and a Macro F1-score of 79.59%. For the IndoBERT model, Focal Loss consistently outperformed both Class Weight (CW) and the combined CW+FL approach, whereas the CW+FL combination tended to reduce performance in several experimental scenarios. The best performance was achieved using a configuration of 10 epochs, a batch size of 32, a learning rate of 2 × 10⁻⁵, and a dropout rate of 0.3, resulting in an Accuracy of 94.32%, Precision of 89.82%, Recall of 91.61%, and a Macro F1-score of 90.58%. The results demonstrate that the LDA–IndoBERT-based ABSA approach outperformed the conventional machine learning models, while Focal Loss proved to be the most effective strategy for handling class imbalance during the fine-tuning process of IndoBERT.
| Item Type: | Thesis (Undergraduate) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | Q Science > QA Mathematics Q Science > QA Mathematics > QA75 Electronic computers. Computer science Q Science > QA Mathematics > QA76.6 Computer Programming T Technology > T Technology (General) |
||||||||||||
| Divisions: | Faculty of Computer Science > Departemen of Information Systems | ||||||||||||
| Depositing User: | Layla Mazidatus Sa`adah | ||||||||||||
| Date Deposited: | 20 Jul 2026 04:40 | ||||||||||||
| Last Modified: | 20 Jul 2026 06:48 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/56160 |
Actions (login required)
![]() |
View Item |
