Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Yicheng, Huang, Zhemin, Yang, Liuxin, Lu, Yumeng, Dai, Zhongdongming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913906879889408
author Fu, Yicheng
Huang, Zhemin
Yang, Liuxin
Lu, Yumeng
Dai, Zhongdongming
author_facet Fu, Yicheng
Huang, Zhemin
Yang, Liuxin
Lu, Yumeng
Dai, Zhongdongming
contents Chinese idioms (Chengyu) are concise four-character expressions steeped in history and culture, whose literal translations often fail to capture their full meaning. This complexity makes them challenging for language models to interpret and use correctly. Existing benchmarks focus on narrow tasks - multiple-choice cloze tests, isolated translation, or simple paraphrasing. We introduce Chengyu-Bench, a comprehensive benchmark featuring three tasks: (1) Evaluative Connotation, classifying idioms as positive or negative; (2) Appropriateness, detecting incorrect idiom usage in context; and (3) Open Cloze, filling blanks in longer passages without options. Chengyu-Bench comprises 2,937 human-verified examples covering 1,765 common idioms sourced from diverse corpora. We evaluate leading LLMs and find they achieve over 95% accuracy on Evaluative Connotation, but only ~85% on Appropriateness and ~40% top-1 accuracy on Open Cloze. Error analysis reveals that most mistakes arise from fundamental misunderstandings of idiom meanings. Chengyu-Bench demonstrates that while LLMs can reliably gauge idiom sentiment, they still struggle to grasp the cultural and contextual nuances essential for proper usage. The benchmark and source code are available at: https://github.com/sofyc/ChengyuBench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
Fu, Yicheng
Huang, Zhemin
Yang, Liuxin
Lu, Yumeng
Dai, Zhongdongming
Computation and Language
Chinese idioms (Chengyu) are concise four-character expressions steeped in history and culture, whose literal translations often fail to capture their full meaning. This complexity makes them challenging for language models to interpret and use correctly. Existing benchmarks focus on narrow tasks - multiple-choice cloze tests, isolated translation, or simple paraphrasing. We introduce Chengyu-Bench, a comprehensive benchmark featuring three tasks: (1) Evaluative Connotation, classifying idioms as positive or negative; (2) Appropriateness, detecting incorrect idiom usage in context; and (3) Open Cloze, filling blanks in longer passages without options. Chengyu-Bench comprises 2,937 human-verified examples covering 1,765 common idioms sourced from diverse corpora. We evaluate leading LLMs and find they achieve over 95% accuracy on Evaluative Connotation, but only ~85% on Appropriateness and ~40% top-1 accuracy on Open Cloze. Error analysis reveals that most mistakes arise from fundamental misunderstandings of idiom meanings. Chengyu-Bench demonstrates that while LLMs can reliably gauge idiom sentiment, they still struggle to grasp the cultural and contextual nuances essential for proper usage. The benchmark and source code are available at: https://github.com/sofyc/ChengyuBench.
title Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
topic Computation and Language
url https://arxiv.org/abs/2506.18105