STACKING ENSEMBLE MODEL WITH MAXIMUM ENTROPY AS META-LEARNER FOR OPINION CLASSIFICATION OF BOOK REVIEWS ON THE GOODREADS PLATFORM

Agustin, Sesillia (2026) STACKING ENSEMBLE MODEL WITH MAXIMUM ENTROPY AS META-LEARNER FOR OPINION CLASSIFICATION OF BOOK REVIEWS ON THE GOODREADS PLATFORM. Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text (Cover)
22083010038.-cover.pdf

Download (1MB)
[img] Text (Bab 1)
22083010038.-bab1.pdf

Download (105kB)
[img] Text (Bab 2)
22083010038.-bab2.pdf
Restricted to Repository staff only until 13 September 2028.

Download (346kB)
[img] Text (Bab 3)
22083010038.-bab3.pdf
Restricted to Repository staff only until 13 September 2028.

Download (395kB)
[img] Text (Bab 4)
22083010038.-bab4.pdf
Restricted to Repository staff only until 13 September 2028.

Download (1MB)
[img] Text (Bab 5)
22083010038.-bab5.pdf

Download (62kB)
[img] Text (Daftar pustaka)
22083010038.-daftarpustaka.pdf

Download (158kB)
[img] Text (Lampiran)
22083010038.-lampiran.pdf
Restricted to Repository staff only

Download (71kB)

Abstract

Book reviews on the Goodreads platform serve as an important source of information for prospective readers in deciding whether a book is worth reading. However, the long formed reviews making it harder for reader to understand the initial meaning of said review, creating a need for an automated sentiment analysis system. This study aims to build a sentiment analysis model to classify Goodreads book reviews into two classes, recommended and not recommended, using a Stacking ensemble approach with Naive Bayes and XGBoost as Base learners, and Maximum Entropy (Logistic Regression) as the meta-learner. The research data consists of 900 reviews scraped from six book titles across different genres, which underwent a preprocessing stage (text cleaning, case folding, tokenizing, stopword removal, and lemmatization) before being numerically represented using the BM25 method. This study also compares two data labeling methods, namely VADER (automatic) and manual annotation, and applies Class Weighting techniques to address class imbalance. The model was tested on three data-splitting ratios (60:40, 70:30, 80:20) for both datasets. The best result was achieved on the manually annotated dataset with a 70:30 ratio, yielding a macro F1-score of 0.76 and a minority-class (not recommended) recall of 0.73. The manually annotated dataset consistently outperformed the VADER-based dataset across all tested ratios, indicating the significant influence of labeling method quality on model performance. The results of this research were implemented into an interactive web-based Dashboard named "The Reading Room," which provides real-time sentiment analysis features on book reviews for users.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorSaputra, Wahyu Syaifullah JauharisNIDN0725088601wahyu.s.j.saputra.if@upnjatim.ac.id
Thesis advisorPratama, Alfan RizaldyNUPTK7938777678130112alfan.fasilkom@upnjatim.ac.id
Subjects: H Social Sciences > HG Finance > HG1709 Data processing
Q Science > QA Mathematics
Q Science > QA Mathematics > QA76.87 Neural computers
Divisions: Faculty of Computer Science > Departemen of Data Science
Depositing User: Sesillia Agustin
Date Deposited: 15 Sep 2026 03:46
Last Modified: 15 Sep 2026 04:00
URI: https://repository.upnjatim.ac.id/id/eprint/60193

Actions (login required)

View Item View Item