Agustin, Sesillia (2026) STACKING ENSEMBLE MODEL WITH MAXIMUM ENTROPY AS META-LEARNER FOR OPINION CLASSIFICATION OF BOOK REVIEWS ON THE GOODREADS PLATFORM. Undergraduate thesis, UPN Veteran Jawa Timur.
|
Text (Cover)
22083010038.-cover.pdf Download (1MB) |
|
|
Text (Bab 1)
22083010038.-bab1.pdf Download (105kB) |
|
|
Text (Bab 2)
22083010038.-bab2.pdf Restricted to Repository staff only until 13 September 2028. Download (346kB) |
|
|
Text (Bab 3)
22083010038.-bab3.pdf Restricted to Repository staff only until 13 September 2028. Download (395kB) |
|
|
Text (Bab 4)
22083010038.-bab4.pdf Restricted to Repository staff only until 13 September 2028. Download (1MB) |
|
|
Text (Bab 5)
22083010038.-bab5.pdf Download (62kB) |
|
|
Text (Daftar pustaka)
22083010038.-daftarpustaka.pdf Download (158kB) |
|
|
Text (Lampiran)
22083010038.-lampiran.pdf Restricted to Repository staff only Download (71kB) |
Abstract
Book reviews on the Goodreads platform serve as an important source of information for prospective readers in deciding whether a book is worth reading. However, the long formed reviews making it harder for reader to understand the initial meaning of said review, creating a need for an automated sentiment analysis system. This study aims to build a sentiment analysis model to classify Goodreads book reviews into two classes, recommended and not recommended, using a Stacking ensemble approach with Naive Bayes and XGBoost as Base learners, and Maximum Entropy (Logistic Regression) as the meta-learner. The research data consists of 900 reviews scraped from six book titles across different genres, which underwent a preprocessing stage (text cleaning, case folding, tokenizing, stopword removal, and lemmatization) before being numerically represented using the BM25 method. This study also compares two data labeling methods, namely VADER (automatic) and manual annotation, and applies Class Weighting techniques to address class imbalance. The model was tested on three data-splitting ratios (60:40, 70:30, 80:20) for both datasets. The best result was achieved on the manually annotated dataset with a 70:30 ratio, yielding a macro F1-score of 0.76 and a minority-class (not recommended) recall of 0.73. The manually annotated dataset consistently outperformed the VADER-based dataset across all tested ratios, indicating the significant influence of labeling method quality on model performance. The results of this research were implemented into an interactive web-based Dashboard named "The Reading Room," which provides real-time sentiment analysis features on book reviews for users.
| Item Type: | Thesis (Undergraduate) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | H Social Sciences > HG Finance > HG1709 Data processing Q Science > QA Mathematics Q Science > QA Mathematics > QA76.87 Neural computers |
||||||||||||
| Divisions: | Faculty of Computer Science > Departemen of Data Science | ||||||||||||
| Depositing User: | Sesillia Agustin | ||||||||||||
| Date Deposited: | 15 Sep 2026 03:46 | ||||||||||||
| Last Modified: | 15 Sep 2026 04:00 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/60193 |
Actions (login required)
![]() |
View Item |
