A Comparison of XGBoost Performance Before and After Hyperparameter Optimization with Optuna for Drinking Water Quality Classification

Sari, Tiara Permata (2026) A Comparison of XGBoost Performance Before and After Hyperparameter Optimization with Optuna for Drinking Water Quality Classification. Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text (Cover)
22081010244.-cover.pdf

Download (2MB)
[img] Text (Bab 1)
22081010244.-bab1.pdf

Download (249kB)
[img] Text (Bab 2)
22081010244.-bab2.pdf
Restricted to Repository staff only until 17 July 2028.

Download (660kB)
[img] Text (Bab 3)
22081010244.-bab3.pdf
Restricted to Repository staff only until 17 July 2028.

Download (929kB)
[img] Text (Bab 4)
22081010244.-bab4.pdf
Restricted to Repository staff only until 17 July 2028.

Download (735kB)
[img] Text (Bab 5)
22081010244.-bab5.pdf

Download (235kB)
[img] Text (Daftar pustaka)
22081010244.-daftarpustaka.pdf

Download (279kB)

Abstract

Drinking water quality that fails to meet potability standards can negatively impact human health; therefore, a data-driven classification approach is necessary to accurately identify water potability. This study evaluates whether hyperparameter tuning can improve the classification performance of the Extreme Gradient Boosting (XGBoost) algorithm for drinking water quality classification while also implementing the resulting model within an interactive web dashboard. The experiments were conducted using the 'Water Quality & Potability' dataset from Kaggle, which contains 3,276 observations, nine predictor variables, and one target variable. Prior to model development, the dataset underwent several preprocessing procedures, including missing-value imputation using Mean Imputation, feature normalization using Min-Max Scaling, class balancing via Oversampling, and dataset partitioning with 80:20, 70:30, and 60:40 train-test split configurations. The baseline XGBoost model was subsequently optimized using Optuna, which applies the Tree-structured Parzen Estimator (TPE) as its sampling strategy for hyperparameter tuning. Model effectiveness was assessed using the Confusion Matrix together with the derived evaluation metrics of accuracy, precision, recall, and F1-score, while an interactive dashboard was developed with Streamlit to visualize the experimental outcomes. Experimental results show that the highest baseline performance was obtained using the 80:20 data split, achieving an accuracy of 0.8438, precision of 0.8647, recall of 0.8150, and an F1-score of 0.8391. After hyperparameter optimization, the optimized XGBoost model produced its best results under the same data partition, with an accuracy of 0.8600, precision of 0.8789, recall of 0.8350, and an F1-score of 0.8564. Compared with the baseline model, these results correspond to improvements of 1.62% in accuracy, 1.42% in precision, 2.00% in recall, and 1.73% in F1-score. The developed Streamlit dashboard effectively visualized the experimental findings through tables, charts, and confusion matrices, thereby supporting a clearer interpretation of model performance. Overall, the findings confirm that hyperparameter tuning with Optuna enhances the predictive capability of the XGBoost model for drinking water quality classification and produces a more effective classifier than the baseline configuration.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorMuttaqin, FaisalNIDN0030058602faisalmuttaqin.if@upnjatim.ac.id
Thesis advisorAnggraeny, Fetty TriNIDN0711028201fettyanggraeny.if@upnjatim.ac.id
Subjects: Q Science > QA Mathematics > QA76.87 Neural computers
Divisions: Faculty of Computer Science
Depositing User: Tiara Permata Sari
Date Deposited: 20 Jul 2026 03:31
Last Modified: 20 Jul 2026 05:28
URI: https://repository.upnjatim.ac.id/id/eprint/55796

Actions (login required)

View Item View Item