Post-Training Quantization Approach for Real-Time Android Face Recognition Based on the MobileFaceNet Model

Hasan, Ferry (2026) Post-Training Quantization Approach for Real-Time Android Face Recognition Based on the MobileFaceNet Model. Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text
22081010085-cover.pdf

Download (484kB)
[img] Text (BAB I)
22081010085-bab1.pdf

Download (247kB)
[img] Text (BAB II)
22081010085-bab2.pdf
Restricted to Repository staff only until 21 July 2029.

Download (993kB)
[img] Text (BAB III)
22081010085-bab3.pdf
Restricted to Repository staff only until 21 July 2029.

Download (603kB)
[img] Text (BAB IV)
22081010085-bab4.pdf
Restricted to Repository staff only until 21 July 2029.

Download (950kB)
[img] Text (BAB V)
22081010085-bab5.pdf

Download (174kB)
[img] Text (Daftar Pustaka)
22081010085-daftarpustaka.pdf

Download (180kB)
[img] Text (Lampiran)
22081010085-lampiran.pdf
Restricted to Repository staff only

Download (428kB)

Abstract

Deep learning-based face recognition systems require high-accuracy models, yet their implementation on mid-to-low-end Android devices is constrained by limited CPU performance, memory capacity, and inference latency. This study aims to optimize a real-time Android-based face recognition system through a post-training quantization approach on a MobileFaceNet model with ArcFace Loss, aiming to achieve the most ideal trade-off among accuracy, inference latency, and resource efficiency on limited hardware. The model was trained on 489,839 images from the CASIA-WebFace dataset over 80 epochs, with early stopping triggered at epoch 72. It was then converted to TensorFlow Lite in three variants: the Baseline Model (FP32), Model A with dynamic range quantization (INT8), and Model B with float16 quantization (FP16). Offline evaluation using a 10-fold cross-validation protocol on the LFW dataset demonstrated stable face verification accuracy ranging from 90.83% to 90.90%, with a maximum absolute accuracy degradation of only 0.07 percentage points across all variants. Model A (INT8) achieved a compression ratio of 3.32×, reducing the size to 1.15 MB, while Model B (FP16) achieved a 1.96× ratio, reducing it to 1.95 MB. On-device evaluation on a Realme 7i powered by a Qualcomm Snapdragon 662 chipset across 27 test sessions showed that all models operated within a total latency range of 210 to 342 milliseconds, with a processing FPS of 3.80 to 4.90. Model B recorded the lowest average latency at 234 ms, which is still 34 ms above the 200 ms industry standard. These findings empirically map the performance gap of mid-to-low-end devices when relying on pure CPU execution. Overall, Model B demonstrated superiority in inference speed and latency stability, whereas Model A proved optimal for model storage efficiency.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorDiyasa, I Gede Susrama MasNIDN0019067008igsusrama.if@upnjatim.ac.id
Thesis advisorJunaidi, AchmadNIDN0710117803achmadjunaidi.if@upnjatim.ac.id
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Divisions: Faculty of Computer Science > Departemen of Informatics
Depositing User: Ferry Hasan
Date Deposited: 22 Jul 2026 03:47
Last Modified: 22 Jul 2026 07:04
URI: https://repository.upnjatim.ac.id/id/eprint/56711

Actions (login required)

View Item View Item