A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yikang, Shen, Yeting, Zhu, Hongao, Xu, Lilong, Qian, Zhiheng, Song, Siyuan, Zhang, Kejia, Tang, Jialong, Zhang, Pei, Yang, Baosong, Wang, Rui, Hu, Hai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908695901765632
author Liu, Yikang
Shen, Yeting
Zhu, Hongao
Xu, Lilong
Qian, Zhiheng
Song, Siyuan
Zhang, Kejia
Tang, Jialong
Zhang, Pei
Yang, Baosong
Wang, Rui
Hu, Hai
author_facet Liu, Yikang
Shen, Yeting
Zhu, Hongao
Xu, Lilong
Qian, Zhiheng
Song, Siyuan
Zhang, Kejia
Tang, Jialong
Zhang, Pei
Yang, Baosong
Wang, Rui
Hu, Hai
contents We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that \textsc{Anaphor}, \textsc{Quantifiers}, and \textsc{Ellipsis} in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06096
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
Liu, Yikang
Shen, Yeting
Zhu, Hongao
Xu, Lilong
Qian, Zhiheng
Song, Siyuan
Zhang, Kejia
Tang, Jialong
Zhang, Pei
Yang, Baosong
Wang, Rui
Hu, Hai
Computation and Language
We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that \textsc{Anaphor}, \textsc{Quantifiers}, and \textsc{Ellipsis} in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs.
title A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
topic Computation and Language
url https://arxiv.org/abs/2411.06096