TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Chenzhuo, Wang, Xinda, Huang, Yue, Lu, Junting, Liu, Ziqian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913979176058880
author Zhao, Chenzhuo
Wang, Xinda
Huang, Yue
Lu, Junting
Liu, Ziqian
author_facet Zhao, Chenzhuo
Wang, Xinda
Huang, Yue
Lu, Junting
Liu, Ziqian
contents While large language models (LLMs) have demonstrated remarkable performance on high-level semantic tasks, they often struggle with fine-grained, token-level understanding and structural reasoning--capabilities that are essential for applications requiring precision and control. We introduce TASE, a comprehensive benchmark designed to evaluate LLMs' ability to perceive and reason about token-level information across languages. TASE covers 10 tasks under two core categories: token awareness and structural understanding, spanning Chinese, English, and Korean, with a 35,927-instance evaluation set and a scalable synthetic data generation pipeline for training. Tasks include character counting, token alignment, syntactic structure parsing, and length constraint satisfaction. We evaluate over 30 leading commercial and open-source LLMs, including O3, Claude 4, Gemini 2.5 Pro, and DeepSeek-R1, and train a custom Qwen2.5-14B model using the GRPO training method. Results show that human performance significantly outpaces current LLMs, revealing persistent weaknesses in token-level reasoning. TASE sheds light on these limitations and provides a new diagnostic lens for future improvements in low-level language understanding and cross-lingual generalization. Our code and dataset are publicly available at https://github.com/cyzcz/Tase .
format Preprint
id arxiv_https___arxiv_org_abs_2508_05468
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
Zhao, Chenzhuo
Wang, Xinda
Huang, Yue
Lu, Junting
Liu, Ziqian
Computation and Language
While large language models (LLMs) have demonstrated remarkable performance on high-level semantic tasks, they often struggle with fine-grained, token-level understanding and structural reasoning--capabilities that are essential for applications requiring precision and control. We introduce TASE, a comprehensive benchmark designed to evaluate LLMs' ability to perceive and reason about token-level information across languages. TASE covers 10 tasks under two core categories: token awareness and structural understanding, spanning Chinese, English, and Korean, with a 35,927-instance evaluation set and a scalable synthetic data generation pipeline for training. Tasks include character counting, token alignment, syntactic structure parsing, and length constraint satisfaction. We evaluate over 30 leading commercial and open-source LLMs, including O3, Claude 4, Gemini 2.5 Pro, and DeepSeek-R1, and train a custom Qwen2.5-14B model using the GRPO training method. Results show that human performance significantly outpaces current LLMs, revealing persistent weaknesses in token-level reasoning. TASE sheds light on these limitations and provides a new diagnostic lens for future improvements in low-level language understanding and cross-lingual generalization. Our code and dataset are publicly available at https://github.com/cyzcz/Tase .
title TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
topic Computation and Language
url https://arxiv.org/abs/2508.05468