Utilization of Complete Blood Count for Anemia Classification Modeling Using CatBoost and SHAP Interpretability Analysis

Rafi Ana, Vika (2026) Utilization of Complete Blood Count for Anemia Classification Modeling Using CatBoost and SHAP Interpretability Analysis. Undergraduate thesis, UPN Veteran JawaTimur.

[img] Text (Cover)
Cover.pdf

Download (291kB)
[img] Text (BAB I)
BAB_1.pdf

Download (100kB)
[img] Text (BAB II)
BAB_2.pdf
Restricted to Repository staff only until 17 July 2028.

Download (802kB)
[img] Text (BAB III)
BAB_3.pdf
Restricted to Repository staff only until 17 July 2028.

Download (825kB)
[img] Text (BAB IV)
BAB_4.pdf
Restricted to Repository staff only until 17 July 2028.

Download (4MB)
[img] Text (BAB V)
BAB_5.pdf

Download (91kB)
[img] Text (Daftar Pustaka)
Daftar Pustaka.pdf

Download (125kB)
[img] Text (Lampiran)
Lampiran.pdf
Restricted to Repository staff only

Download (357kB)

Abstract

This study proposes an anemia classification model based on Complete Blood Count (CBC) data, as hematological parameters exhibit distinct patterns across different diagnosis classes. The input variables consist of CBC parameters, including HGB, RBC, MCV, MCH, MCHC, WBC, and PLT. The classification process was performed using the CatBoost algorithm, while SHAP was employed to explain the contribution of each CBC feature to the model predictions. The dataset utilized in this research was the Anemia Types Classification dataset obtained from Kaggle. Initially, the dataset comprised 1,281 samples with 14 CBC attributes and 9 diagnosis classes. Following the removal of duplicate records, the final dataset contained 1,232 samples. The cleaned dataset was partitioned into training and testing subsets using an 80:20 stratified split to preserve the class distribution. Three experimental settings were investigated, consisting of the baseline CatBoost model, CatBoost with class weighting, and CatBoost with both class weighting and hyperparameter tuning. Among these configurations, the CatBoost model with class weighting achieved the highest performance, obtaining Accuracy, Macro F1-score, and Weighted F1-score values of 1.0000 on the testing dataset. In contrast, applying hyperparameter tuning did not improve the predictive performance and instead resulted in a considerably longer training time. The SHAP analysis further revealed that HGB, MCV, MCH, MCHC, WBC, PLT, and RBC were the most influential CBC features contributing to the classification process. Overall, the findings demonstrate that incorporating class weighting enhanced the classification capability of CatBoost, whereas SHAP improved the interpretability of the model by explaining how individual CBC features influenced the prediction outcomes.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorYulia Puspaningrum, EvaNIDN0005078908evapuspaningrum.if@upnjatim.ac.id
Thesis advisorMumpuni, RetnoNIDN0016078703retnomumpuni.if@upnjatim.ac.id
Subjects: T Technology > T Technology (General)
Divisions: Faculty of Computer Science > Departemen of Informatics
Depositing User: Vika Rafi Ana
Date Deposited: 17 Jul 2026 07:48
Last Modified: 17 Jul 2026 08:22
URI: https://repository.upnjatim.ac.id/id/eprint/55917

Actions (login required)

View Item View Item