Saved in:
Bibliographic Details
Main Authors: Zeng, Jiayi, Feng, Yizhe, He, Mengliang, Lei, Wenhui, Zhang, Wei, Liu, Zeming, Shi, Xiaoming, Zhou, Aimin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.00064
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918041082658816
author Zeng, Jiayi
Feng, Yizhe
He, Mengliang
Lei, Wenhui
Zhang, Wei
Liu, Zeming
Shi, Xiaoming
Zhou, Aimin
author_facet Zeng, Jiayi
Feng, Yizhe
He, Mengliang
Lei, Wenhui
Zhang, Wei
Liu, Zeming
Shi, Xiaoming
Zhou, Aimin
contents Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. The dataset will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
Zeng, Jiayi
Feng, Yizhe
He, Mengliang
Lei, Wenhui
Zhang, Wei
Liu, Zeming
Shi, Xiaoming
Zhou, Aimin
Computation and Language
Artificial Intelligence
Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. The dataset will be publicly available.
title Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.00064