Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sasaki, Mutsumi, Kamoda, Go, Takahashi, Ryosuke, Sato, Kosuke, Inui, Kentaro, Sakaguchi, Keisuke, Heinzerling, Benjamin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912703233130496
author Sasaki, Mutsumi
Kamoda, Go
Takahashi, Ryosuke
Sato, Kosuke
Inui, Kentaro
Sakaguchi, Keisuke
Heinzerling, Benjamin
author_facet Sasaki, Mutsumi
Kamoda, Go
Takahashi, Ryosuke
Sato, Kosuke
Inui, Kentaro
Sakaguchi, Keisuke
Heinzerling, Benjamin
contents Temporal reasoning and knowledge are essential capabilities for language models (LMs). While much prior work has analyzed and improved temporal reasoning in LMs, most studies have focused solely on the Gregorian calendar. However, many non-Gregorian systems, such as the Japanese, Hijri, and Hebrew calendars, are in active use and reflect culturally grounded conceptions of time. If and how well current LMs can accurately handle such non-Gregorian calendars has not been evaluated so far. Here, we present a systematic evaluation of how well language models handle one such non-Gregorian system: the Japanese wareki. We create datasets that require temporal knowledge and reasoning in using wareki dates. Evaluating open and closed LMs, we find that some models can perform calendar conversions, but GPT-4o, Deepseek V3, and even Japanese-centric models struggle with Japanese calendar arithmetic and knowledge involving wareki dates. Error analysis suggests corpus frequency of Japanese calendar expressions and a Gregorian bias in the model's knowledge as possible explanations. Our results show the importance of developing LMs that are better equipped for culture-specific tasks such as calendar understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
Sasaki, Mutsumi
Kamoda, Go
Takahashi, Ryosuke
Sato, Kosuke
Inui, Kentaro
Sakaguchi, Keisuke
Heinzerling, Benjamin
Computation and Language
Temporal reasoning and knowledge are essential capabilities for language models (LMs). While much prior work has analyzed and improved temporal reasoning in LMs, most studies have focused solely on the Gregorian calendar. However, many non-Gregorian systems, such as the Japanese, Hijri, and Hebrew calendars, are in active use and reflect culturally grounded conceptions of time. If and how well current LMs can accurately handle such non-Gregorian calendars has not been evaluated so far. Here, we present a systematic evaluation of how well language models handle one such non-Gregorian system: the Japanese wareki. We create datasets that require temporal knowledge and reasoning in using wareki dates. Evaluating open and closed LMs, we find that some models can perform calendar conversions, but GPT-4o, Deepseek V3, and even Japanese-centric models struggle with Japanese calendar arithmetic and knowledge involving wareki dates. Error analysis suggests corpus frequency of Japanese calendar expressions and a Gregorian bias in the model's knowledge as possible explanations. Our results show the importance of developing LMs that are better equipped for culture-specific tasks such as calendar understanding.
title Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
topic Computation and Language
url https://arxiv.org/abs/2509.04432