Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Wenhan, Hu, Tianyi, Zheng, Jingyi, Sun, Zhen, Zhao, Yuemeng, Liu, Yule, He, Xinlei, Huang, Xinyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918038863872000
author Dong, Wenhan
Hu, Tianyi
Zheng, Jingyi
Sun, Zhen
Zhao, Yuemeng
Liu, Yule
He, Xinlei
Huang, Xinyi
author_facet Dong, Wenhan
Hu, Tianyi
Zheng, Jingyi
Sun, Zhen
Zhao, Yuemeng
Liu, Yule
He, Xinlei
Huang, Xinyi
contents Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs). Currently, research primarily relies on multiple benchmarks and automated evaluation metrics to assess these abilities. However, our study reveals novel insights into the limitations of existing methods, as they often yield misleading results that fail to uncover key issues, such as shortcut-taking behaviors, rigid patterns, and premature task termination. These issues obscure the true reasoning capabilities of LLMs and undermine the reliability of evaluations. To address these limitations, we propose a refined set of evaluation standards, including inspection of reasoning paths, diversified assessment metrics, and comparative analyses with human performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks
Dong, Wenhan
Hu, Tianyi
Zheng, Jingyi
Sun, Zhen
Zhao, Yuemeng
Liu, Yule
He, Xinlei
Huang, Xinyi
Computation and Language
Machine Learning
Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs). Currently, research primarily relies on multiple benchmarks and automated evaluation metrics to assess these abilities. However, our study reveals novel insights into the limitations of existing methods, as they often yield misleading results that fail to uncover key issues, such as shortcut-taking behaviors, rigid patterns, and premature task termination. These issues obscure the true reasoning capabilities of LLMs and undermine the reliability of evaluations. To address these limitations, we propose a refined set of evaluation standards, including inspection of reasoning paths, diversified assessment metrics, and comparative analyses with human performance.
title Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.23843