Do Language Models Understand Honorific Systems in Javanese?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farhansyah, Mohammad Rifqi, Darmawan, Iwan, Kusumawardhana, Adryan, Winata, Genta Indra, Aji, Alham Fikri, Wijaya, Derry Tanti
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915312910204928
author Farhansyah, Mohammad Rifqi
Darmawan, Iwan
Kusumawardhana, Adryan
Winata, Genta Indra
Aji, Alham Fikri
Wijaya, Derry Tanti
author_facet Farhansyah, Mohammad Rifqi
Darmawan, Iwan
Kusumawardhana, Adryan
Winata, Genta Indra
Aji, Alham Fikri
Wijaya, Derry Tanti
contents The Javanese language features a complex system of honorifics that vary according to the social status of the speaker, listener, and referent. Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natural language processing (NLP) tasks. In this paper, we present Unggah-Ungguh, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Using Unggah-Ungguh, we assess the ability of language models (LMs) to process various levels of Javanese honorifics through classification and machine translation tasks. To further evaluate cross-lingual LMs, we conduct machine translation experiments between Javanese (at specific honorific levels) and Indonesian. Additionally, we explore whether LMs can generate contextually appropriate Javanese honorifics in conversation tasks, where the honorific usage should align with the social role and contextual cues. Our findings indicate that current LMs struggle with most honorific levels, exhibitinga bias toward certain honorific tiers.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Language Models Understand Honorific Systems in Javanese?
Farhansyah, Mohammad Rifqi
Darmawan, Iwan
Kusumawardhana, Adryan
Winata, Genta Indra
Aji, Alham Fikri
Wijaya, Derry Tanti
Computation and Language
The Javanese language features a complex system of honorifics that vary according to the social status of the speaker, listener, and referent. Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natural language processing (NLP) tasks. In this paper, we present Unggah-Ungguh, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Using Unggah-Ungguh, we assess the ability of language models (LMs) to process various levels of Javanese honorifics through classification and machine translation tasks. To further evaluate cross-lingual LMs, we conduct machine translation experiments between Javanese (at specific honorific levels) and Indonesian. Additionally, we explore whether LMs can generate contextually appropriate Javanese honorifics in conversation tasks, where the honorific usage should align with the social role and contextual cues. Our findings indicate that current LMs struggle with most honorific levels, exhibitinga bias toward certain honorific tiers.
title Do Language Models Understand Honorific Systems in Javanese?
topic Computation and Language
url https://arxiv.org/abs/2502.20864