Analisis Struktur Semantik Teks Kebudayaan Menggunakan Transformer-Based Embedding dan Clustering (Studi Kasus: Teks Dewi Durga)

Widyadhana, Nala (2026) Analisis Struktur Semantik Teks Kebudayaan Menggunakan Transformer-Based Embedding dan Clustering (Studi Kasus: Teks Dewi Durga). Undergraduate thesis, UPN Veteran Jawa Timur.

[img] Text (COVER)
22082010195-Cover.pdf

Download (2MB)
[img] Text (BAB 1)
22082010195-Bab 1.pdf

Download (158kB)
[img] Text (BAB 2)
22082010195-Bab 2.pdf
Restricted to Repository staff only until 22 July 2029.

Download (626kB)
[img] Text (BAB 3)
22082010195-Bab 3.pdf
Restricted to Repository staff only until 27 July 2029.

Download (475kB)
[img] Text (BAB 4)
22082010195-Bab 4.pdf
Restricted to Repository staff only until 22 July 2029.

Download (1MB)
[img] Text (BAB 5)
22082010195-Bab 5.pdf

Download (120kB)
[img] Text (DAFTAR PUSTAKA)
22082010195-Daftar Pustaka.pdf

Download (233kB)
[img] Text (LAMPIRAN)
22082010195-Lampiran.pdf

Download (551kB)

Abstract

Cultural texts contain complex semantic relations because they include historical, symbolic, spiritual, and socio-cultural elements. This study aims to analyze the semantic structure of Dewi Durga cultural texts and evaluate the quality and consistency of semantic representations generated by several pretrained Sentence Transformer models through embedding parameter scenarios. This research applies a text mining approach using 1,620 valid data units structured in the Context–Question–Answer (CQA) format. The five pretrained models used in this study are all-MiniLM-L6-v2, paraphrase-mpnet-base-v2, all-distilroberta-v1, paraphrase-multilingual-mpnet-base-v2, and distiluse-base-multilingual-cased-v2. The experiment was conducted through 30 scenarios combining model variations, embedding normalization, and PCA-based dimensionality reduction. Text clustering was performed using K-Means with k=5, while clustering quality was evaluated using Silhouette Score, Davies–Bouldin Index, and Calinski–Harabasz Index, integrated through the Multi-Metric Ranking Framework. The results show that Dewi Durga texts can be mapped into five semantic themes: historical representation in Javanese society, ritual and spirituality of worship, spiritual dimensions and socio-cultural interaction, academic literature and Hindu tradition studies, and worship practices and cultural artifacts. Scenario evaluation shows that S17, namely all-distilroberta-v1 with active normalization and PCA 50, is the most aggregatively consistent configuration at k=5, with a Silhouette Score of 0.090608, Davies–Bouldin Index of 2.888233, and Calinski–Harabasz Index of 131.543594. These findings indicate that semantic embedding can identify cultural meaning similarity despite different linguistic variations.

Item Type: Thesis (Undergraduate)
Contributors:
ContributionContributorsNIDN/NIDKEmail
Thesis advisorWibowo, Nur CahyoNIDN0717037901nurcahyo.si@upnjatim.ac.id
Thesis advisorSuryanto, Tri Lathif MardiNIDN0025028902trilathif.si@upnjatim.ac.id
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5105 Computer Network
T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5105.882 Internet
Divisions: Faculty of Computer Science > Departemen of Information Systems
Depositing User: Nala Widyadhana
Date Deposited: 22 Jul 2026 07:58
Last Modified: 22 Jul 2026 08:44
URI: https://repository.upnjatim.ac.id/id/eprint/57543

Actions (login required)

View Item View Item