Alam, Fajar Indra Nur (2026) VALIDASI BIODATA MAHASISWA MELALUI PENDEKATAN NLP BERBASIS INDOBERT. Masters thesis, Universitas Pembangunan Nasional Veteran Jawa Timur.
|
Text (Cover)
Cover Tesis.pdf Download (1MB) |
|
|
Text (Bab 1)
Bab 1.pdf Download (204kB) |
|
|
Text (Bab 2)
Bab 2.pdf Restricted to Repository staff only until 10 August 2029. Download (1MB) |
|
|
Text (Bab 3)
Bab 3.pdf Restricted to Repository staff only until 10 August 2029. Download (1MB) |
|
|
Text (Bab 4)
Bab 4.pdf Restricted to Repository staff only until 10 August 2029. Download (1MB) |
|
|
Text (Bab 5)
Bab 5.pdf Download (437kB) |
|
|
Text (Daftar Pustaka)
Daftar Pustaka.pdf Download (319kB) |
|
|
Text (Lampiran)
Lampiran.pdf Restricted to Repository staff only until 10 August 2029. Download (333kB) |
Abstract
The bachelor's degree graduation registration process in higher education is often hindered by discrepancies in student biodata between the academic database (SIAMIK) and the physical high school diploma document. Manual verification is susceptible to human error and is time-consuming. This research aims to develop an intelligent system to automatically validate graduation registration biodata to improve the efficiency and accuracy of academic administration. The proposed method combines three continuous approaches: text extraction from scanned diplomas using a Vision-Language Model interface (Gemini API), Named Entity Recognition (NER) utilizing the IndoBERT language model fine-tuned on 1,500 CoNLL-formatted data corpora, and string value matching through a Fuzzy String Matching algorithm (Token Set Ratio). System performance evaluation is measured through the Character Error Rate (CER) metric for text acquisition, classification evaluation (Accuracy, Precision, Recall, and F1-Score) for NER, and the validity accuracy at final validation. The test results show that the visual extraction stage achieves an excellent average CER of 4.2%. The accuracy result reached 99%. In the semantic understanding stage, the IndoBERT architecture demonstrates outstanding performance in extracting target entities (Student Name, Place of Birth, and Date of Birth) with a testing F1-Score reaching 90%. Furthermore, the tolerance matching module (Fuzzy Matching) is proven capable of handling text writing variations to precisely validate the correctness of diploma data against the SIAMIK reference based on a 90% similarity threshold. In conclusion, the integration of the IndoBERT-based natural language processing approach significantly succeeds in minimizing manual intervention, automating document authorization, and optimizing bachelor's degree graduation registration services.
| Item Type: | Thesis (Masters) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | T Technology > T Technology (General) T Technology > T Technology (General) > T58.6-58.62 Management Information Systems |
||||||||||||
| Divisions: | Faculty of Computer Science > Magister Information Technology | ||||||||||||
| Depositing User: | Fajar Indra Nur Alam | ||||||||||||
| Date Deposited: | 10 Aug 2026 07:45 | ||||||||||||
| Last Modified: | 10 Aug 2026 07:45 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/58476 |
Actions (login required)
![]() |
View Item |
