Widyadhana, Nala (2026) Analisis Struktur Semantik Teks Kebudayaan Menggunakan Transformer-Based Embedding dan Clustering (Studi Kasus: Teks Dewi Durga). Undergraduate thesis, UPN Veteran Jawa Timur.
|
Text (COVER)
22082010195-Cover.pdf Download (2MB) |
|
|
Text (BAB 1)
22082010195-Bab 1.pdf Download (158kB) |
|
|
Text (BAB 2)
22082010195-Bab 2.pdf Restricted to Repository staff only until 22 July 2029. Download (626kB) |
|
|
Text (BAB 3)
22082010195-Bab 3.pdf Restricted to Repository staff only until 27 July 2029. Download (475kB) |
|
|
Text (BAB 4)
22082010195-Bab 4.pdf Restricted to Repository staff only until 22 July 2029. Download (1MB) |
|
|
Text (BAB 5)
22082010195-Bab 5.pdf Download (120kB) |
|
|
Text (DAFTAR PUSTAKA)
22082010195-Daftar Pustaka.pdf Download (233kB) |
|
|
Text (LAMPIRAN)
22082010195-Lampiran.pdf Download (551kB) |
Abstract
Cultural texts contain complex semantic relations because they include historical, symbolic, spiritual, and socio-cultural elements. This study aims to analyze the semantic structure of Dewi Durga cultural texts and evaluate the quality and consistency of semantic representations generated by several pretrained Sentence Transformer models through embedding parameter scenarios. This research applies a text mining approach using 1,620 valid data units structured in the Context–Question–Answer (CQA) format. The five pretrained models used in this study are all-MiniLM-L6-v2, paraphrase-mpnet-base-v2, all-distilroberta-v1, paraphrase-multilingual-mpnet-base-v2, and distiluse-base-multilingual-cased-v2. The experiment was conducted through 30 scenarios combining model variations, embedding normalization, and PCA-based dimensionality reduction. Text clustering was performed using K-Means with k=5, while clustering quality was evaluated using Silhouette Score, Davies–Bouldin Index, and Calinski–Harabasz Index, integrated through the Multi-Metric Ranking Framework. The results show that Dewi Durga texts can be mapped into five semantic themes: historical representation in Javanese society, ritual and spirituality of worship, spiritual dimensions and socio-cultural interaction, academic literature and Hindu tradition studies, and worship practices and cultural artifacts. Scenario evaluation shows that S17, namely all-distilroberta-v1 with active normalization and PCA 50, is the most aggregatively consistent configuration at k=5, with a Silhouette Score of 0.090608, Davies–Bouldin Index of 2.888233, and Calinski–Harabasz Index of 131.543594. These findings indicate that semantic embedding can identify cultural meaning similarity despite different linguistic variations.
| Item Type: | Thesis (Undergraduate) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
||||||||||||
| Subjects: | T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5105 Computer Network T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5105.882 Internet |
||||||||||||
| Divisions: | Faculty of Computer Science > Departemen of Information Systems | ||||||||||||
| Depositing User: | Nala Widyadhana | ||||||||||||
| Date Deposited: | 22 Jul 2026 07:58 | ||||||||||||
| Last Modified: | 22 Jul 2026 08:44 | ||||||||||||
| URI: | https://repository.upnjatim.ac.id/id/eprint/57543 |
Actions (login required)
![]() |
View Item |
