Rafi Ana, Vika (2026) Utilization of Complete Blood Count for Anemia Classification Modeling Using CatBoost and SHAP Interpretability Analysis. Undergraduate thesis, UPN Veteran JawaTimur.
|
Text (Cover)
Cover.pdf Download (291kB) |
|
|
Text (BAB I)
BAB_1.pdf Download (100kB) |
|
|
Text (BAB II)
BAB_2.pdf Restricted to Repository staff only until 17 July 2028. Download (802kB) |
|
|
Text (BAB III)
BAB_3.pdf Restricted to Repository staff only until 17 July 2028. Download (825kB) |
|
|
Text (BAB IV)
BAB_4.pdf Restricted to Repository staff only until 17 July 2028. Download (4MB) |
|
|
Text (BAB V)
BAB_5.pdf Download (91kB) |
|
|
Text (Daftar Pustaka)
Daftar Pustaka.pdf Download (125kB) |
|
|
Text (Lampiran)
Lampiran.pdf Restricted to Repository staff only Download (357kB) |
Abstract
This study proposes an anemia classification model based on Complete Blood Count (CBC) data, as hematological parameters exhibit distinct patterns across different diagnosis classes. The input variables consist of CBC parameters, including HGB, RBC, MCV, MCH, MCHC, WBC, and PLT. The classification process was performed using the CatBoost algorithm, while SHAP was employed to explain the contribution of each CBC feature to the model predictions. The dataset utilized in this research was the Anemia Types Classification dataset obtained from Kaggle. Initially, the dataset comprised 1,281 samples with 14 CBC attributes and 9 diagnosis classes. Following the removal of duplicate records, the final dataset contained 1,232 samples. The cleaned dataset was partitioned into training and testing subsets using an 80:20 stratified split to preserve the class distribution. Three experimental settings were investigated, consisting of the baseline CatBoost model, CatBoost with class weighting, and CatBoost with both class weighting and hyperparameter tuning. Among these configurations, the CatBoost model with class weighting achieved the highest performance, obtaining Accuracy, Macro F1-score, and Weighted F1-score values of 1.0000 on the testing dataset. In contrast, applying hyperparameter tuning did not improve the predictive performance and instead resulted in a considerably longer training time. The SHAP analysis further revealed that HGB, MCV, MCH, MCHC, WBC, PLT, and RBC were the most influential CBC features contributing to the classification process. Overall, the findings demonstrate that incorporating class weighting enhanced the classification capability of CatBoost, whereas SHAP improved the interpretability of the model by explaining how individual CBC features influenced the prediction outcomes.
| Item Type: | Thesis (Undergraduate) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | T Technology > T Technology (General) | ||||||||||||
| Divisions: | Faculty of Computer Science > Departemen of Informatics | ||||||||||||
| Depositing User: | Vika Rafi Ana | ||||||||||||
| Date Deposited: | 17 Jul 2026 07:48 | ||||||||||||
| Last Modified: | 17 Jul 2026 08:22 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/55917 |
Actions (login required)
![]() |
View Item |
