Saved in:
Bibliographic Details
Main Authors: Li, Yifei, Chen, Guanyi, He, Tingting
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.31056
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910271936659456
author Li, Yifei
Chen, Guanyi
He, Tingting
author_facet Li, Yifei
Chen, Guanyi
He, Tingting
contents Zero Pronouns (ZPs) are a pervasive linguistic phenomenon in pro-drop languages such as Chinese and have long posed a challenge for natural language processing systems. Although Large Language Models (LLMs) perform well on many Chinese language tasks, their ability to process ZPs remains poorly understood. We conduct a systematic investigation of LLMs' handling of Chinese ZPs through a sequence of linguistically motivated tasks, including identification, referentiality classification, referential type classification, resolution, and translation. A diverse set of LLMs is evaluated across all tasks. Our results show that Chinese ZPs remain highly challenging for current LLMs, particularly for upstream tasks such as identification and referentiality classification. Performance on downstream tasks, such as ZP translation, is also consistently low: even state-of-the-art reasoning-oriented LLMs correctly translate fewer than half of Chinese ZPs into English.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31056
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Much Do LLMs Know About Chinese Zero Pronouns?
Li, Yifei
Chen, Guanyi
He, Tingting
Computation and Language
Zero Pronouns (ZPs) are a pervasive linguistic phenomenon in pro-drop languages such as Chinese and have long posed a challenge for natural language processing systems. Although Large Language Models (LLMs) perform well on many Chinese language tasks, their ability to process ZPs remains poorly understood. We conduct a systematic investigation of LLMs' handling of Chinese ZPs through a sequence of linguistically motivated tasks, including identification, referentiality classification, referential type classification, resolution, and translation. A diverse set of LLMs is evaluated across all tasks. Our results show that Chinese ZPs remain highly challenging for current LLMs, particularly for upstream tasks such as identification and referentiality classification. Performance on downstream tasks, such as ZP translation, is also consistently low: even state-of-the-art reasoning-oriented LLMs correctly translate fewer than half of Chinese ZPs into English.
title How Much Do LLMs Know About Chinese Zero Pronouns?
topic Computation and Language
url https://arxiv.org/abs/2605.31056