Basuki, Athiyah Fitriyani (2026) Application of the Cross Vision Transformer Method for the Classification of AI Generated or Manual Illustrator Images. Undergraduate thesis, UPN Veteran Jawa Timur.
|
Text (Cover)
20081010056.cover.pdf Download (489kB) |
|
|
Text (Bab 1)
20081010056.-bab1.pdf Download (292kB) |
|
|
Text (Bab 2)
20081010056.-bab2.pdf Restricted to Repository staff only until 17 July 2029. Download (665kB) |
|
|
Text (Bab 3)
20081010056.-bab3.pdf Restricted to Repository staff only until 17 July 2029. Download (826kB) |
|
|
Text (Bab 4)
20081010056.-bab4.pdf Restricted to Repository staff only until 17 July 2029. Download (875kB) |
|
|
Text (Bab 5)
20081010056.-bab5.pdf Download (419kB) |
|
|
Text (Daftar Pustaka)
20081010056.-daftarpustaka.pdf Download (331kB) |
|
|
Text (Lampiran)
20081010056.-lampiran.pdf Restricted to Repository staff only Download (502kB) |
Abstract
This study aims to develop an AI-Generated image classification system and Manual Illustrator images using the Cross Vision Transformer (CrossViT) method as the main architecture with a dual-branch multi-scale mechanism. The dataset used consists of 201 images divided into 100 AI-generated images generated by Gemini's generative model, as well as 101 Illustrator Manual images sourced from the ARIA Dataset. All images have gone through the preprocessing stage in the form of resize to 256×256 pixels, center crop to 224×224 pixels, and pixel normalization using ImageNet statistics. The CrossViT-15 architecture implemented uses two parallel processing branches, namely S-Branch with patch 16×16 generates 196 tokens to capture fine details, and L-Branch with patch 32×32 generates 49 tokens to understand the global context. The two branches are connected through a CLS token-based cross-attention mechanism that allows for efficient cross-scale information exchange. The model was trained using the AdamW optimizer with a learning rate of 1×10⁻⁴, weight decay 0.35, dropout 0.6, and WarmupCosine Scheduler for 5 epochs until it reached an early stopping condition. The results showed that the CrossViT model managed to achieve a validation accuracy of 90.00% and a test accuracy of 93.55% with a Precision value of 88.24%, Recall of 100.00%, and an F1-Score of 93.75%. A Recall value of 100% indicates that all AI-Generated images were successfully detected without any of them passing as manual images. From the test of five hyperparameter scenarios, Scenario 4 with a learning rate of 1×10⁻³ and a batch size of 64 produced the best performance with a test accuracy of 96.67% and an F1-Score of 96.77%. The model is then deployed using the Streamlite framework as an interactive web application that supports single and batch image predictions. This system is expected to be a technical foundation for digital platforms and copyright institutions in verifying the authenticity of works of art in the era of generative AI development.
| Item Type: | Thesis (Undergraduate) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | Q Science > QA Mathematics > QA76.87 Neural computers | ||||||||||||
| Divisions: | Faculty of Computer Science > Departemen of Informatics | ||||||||||||
| Depositing User: | Athiyah Fitriyani Basuki | ||||||||||||
| Date Deposited: | 20 Jul 2026 02:00 | ||||||||||||
| Last Modified: | 20 Jul 2026 02:00 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/55861 |
Actions (login required)
![]() |
View Item |
