Klasifikasi Ulasan Palsu Pada Website Female Daily Menggunakan Algoritma Random Forest

Halizah, Nur (2026) Klasifikasi Ulasan Palsu Pada Website Female Daily Menggunakan Algoritma Random Forest. Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text (Cover)
21082010128-cover.pdf

Download (600kB)
[img] Text (Bab 1)
21082010128-bab1.pdf

Download (514kB)
[img] Text (Bab 2)
21082010128-bab2.pdf
Restricted to Repository staff only until 9 September 2028.

Download (592kB)
[img] Text (Bab 3)
21082010128-bab3.pdf
Restricted to Repository staff only until 9 September 2028.

Download (246kB)
[img] Text (Bab 4)
21082010128-bab4.pdf
Restricted to Repository staff only until 9 September 2028.

Download (767kB)
[img] Text (Bab 5)
21082010128-bab5.pdf

Download (393kB)
[img] Text (Daftar Pustaka)
21082010128-daftarpustaka.pdf

Download (321kB)
[img] Text (Lampiran)
21082010128-lampiran.pdf
Restricted to Repository staff only until 9 September 2028.

Download (612kB)

Abstract

Fake reviews can reduce the objectivity of information used by consumers in purchase decisions. This study develops a fake-review classifier for the Female Daily website using Random Forest with semi-supervised pseudo-labeling based on confidence and uncertainty. The dataset contains 5,724 reviews of Implora Day to Day Series products. A total of 600 reviews were manually labeled by three annotators, consisting of 176 authentic and 424 fake reviews, with a Fleiss' Kappa of 0.815591; the remaining 5,124 reviews were treated as unlabeled data. The labeled data were evaluated using 70:30 and 60:40 train-validation scenarios for model development and selection. FastText achieved the highest mean validation Macro-F1 among the text representations at 0.8071, outperforming TF-IDF and Transformer/Sentence Embedding. FastText was then fused with seven non-text features: Stars, Profile Age, Usage Period, Recommend, Product Name, Product Shade, and Purchase Point. Based on mean validation Macro-F1, Uncertainty Pseudo-Label ranked first at 0.8126, followed by Standard Pseudo-Label at 0.8103, Supervised Random Forest at 0.8041, and Confidence and Confidence+Uncertainty at 0.7972. The selected model was subsequently evaluated on manually labeled external datasets to assess generalization.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorArifiyanti, Amalia AnjaniNIDN0712089201amalia_anjani.fik@upnjatim.ac.id
Thesis advisorSugata, Tri Luhur IndayantiNIDN8948770671230422tri.luhur.fasilkom@upnjatim.ac.id
Subjects: T Technology > T Technology (General)
Divisions: Faculty of Computer Science > Departemen of Information Systems
Depositing User: Liza Nur Halizah
Date Deposited: 09 Sep 2026 08:05
Last Modified: 09 Sep 2026 08:05
URI: https://repository.upnjatim.ac.id/id/eprint/59994

Actions (login required)

View Item View Item