Development of a Real-Time IoT-Based BISINDO Sign Language Translation System Using YOLO

Widison, Daffin Tanjiro (2026) Development of a Real-Time IoT-Based BISINDO Sign Language Translation System Using YOLO. Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text (Cover)
22083010010-cover.pdf

Download (767kB)
[img] Text (Bab 1)
22083010010-bab1.pdf

Download (210kB)
[img] Text (Bab 2)
22083010010-bab2.pdf
Restricted to Repository staff only until 20 July 2029.

Download (1MB)
[img] Text (Bab 3)
22083010010-bab3.pdf
Restricted to Repository staff only until 20 July 2029.

Download (1MB)
[img] Text (Bab 4)
22083010010-bab4.pdf
Restricted to Repository staff only until 20 July 2029.

Download (3MB)
[img] Text (Bab 5)
22083010010-bab5.pdf

Download (152kB)
[img] Text (Daftar Pustaka)
22083010010-daftarpustaka.pdf

Download (186kB)
[img] Text (Lampiran)
22083010010-lampiran.pdf
Restricted to Repository staff only until 20 July 2029.

Download (204kB)

Abstract

Indonesian Sign Language (BISINDO) is the primary means of communication for the deaf, but the general public’s limited understanding of BISINDO remains a barrier to everyday communication. This study aims to develop an Internet of Things (IoT)-based BISINDO translation system capable of detecting and translating gestures in real time using a deep learning-based object detection method. This study conducted a comparative analysis of seven object detection models—YOLOv8m, YOLOv9m, YOLOv10m, YOLO11m, YOLO26m, Faster RCNN, and SSD—using a combined dataset comprising 7,169 primary and secondary images that include alphabet gestures and 30 basic BISINDO vocabulary words. The training and evaluation processes were conducted in a computing environment using an NVIDIA GeForce GTX 1050 GPU to determine the model with the best balance between accuracy and inference speed. The selected model was subsequently implemented on an ESP32-CAM-based IoT system integrated with a Flask-based inference server. Test results show that YOLOv8m was selected as the best model because it offers the most optimal balance between accuracy and inference speed, with an mAP50 of 98.0%, an mAP50-95 of 79.7%, and an inference latency of 53.64 ms. Implementing the model in the system yielded an accuracy of 95.4%, a precision of 97.9%, a recall of 95.4%, and an F1-score of 95.7% on recorded data obtained directly through the IoT prototype. The research results show that the combination of computer vision technology, the YOLO model, and IoT can produce an accurate and responsive BISINDO translation system with the potential to support communication between the deaf community and the general public.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorSaputra, Wahyu Syaifullah JauharisNIDN0725088601wahyu.s.j.saputra.if@upnjatim.ac.id
Thesis advisorPratama, Alfan RizaldyNUPTK7938777678130112alfan.fasilkom@upnjatim.ac.id
Subjects: Q Science > QA Mathematics > QA76.6 Computer Programming
Divisions: Faculty of Computer Science > Departemen of Data Science
Depositing User: Daffin Daffin Widison
Date Deposited: 21 Jul 2026 01:14
Last Modified: 21 Jul 2026 01:59
URI: https://repository.upnjatim.ac.id/id/eprint/56307

Actions (login required)

View Item View Item