HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Chao, Yang, Wenkui, Sun, Hao, Liao, Yuqi, Jiang, Qianyi, Zhou, Kai, Cao, Jie, He, Ran, Huang, Huaibo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908978063081472
author Jin, Chao
Yang, Wenkui
Sun, Hao
Liao, Yuqi
Jiang, Qianyi
Zhou, Kai
Cao, Jie
He, Ran
Huang, Huaibo
author_facet Jin, Chao
Yang, Wenkui
Sun, Hao
Liao, Yuqi
Jiang, Qianyi
Zhou, Kai
Cao, Jie
He, Ran
Huang, Huaibo
contents While progress in GUI agents has been largely driven by industrial-scale training, ungrounded hallucinations often trigger cascading failures in real-world deployments.Unlike general VLM domains, the GUI agent field lacks a hallucination-focused suite for fine-grained diagnosis, reliable evaluation, and targeted mitigation.To bridge this gap, we introduce HalluClear, a comprehensive suite for hallucination mitigation in GUI agents as a complement to computation-intensive scaling. HalluClear comprises: (1) a GUI-specific hallucination taxonomy derived from empirical failure analysis; (2) a calibrated three-stage evaluation workflow which enhances VLM-as-a-judge reliability via expert-annotated benchmarking and ensemble credibility estimation; and (3) a mitigation scheme based on closed-loop structured reasoning, enabling lightweight continual post-training with cold-start initialization for both generalist and GUI-specialist agents. Experiments across representative agents and public benchmarks demonstrate that post-training on only 9K samples within our suite can significantly reduce hallucinations, thereby improving grounding and action fidelity, offering a compute-efficient pathway to robust GUI automation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17284
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
Jin, Chao
Yang, Wenkui
Sun, Hao
Liao, Yuqi
Jiang, Qianyi
Zhou, Kai
Cao, Jie
He, Ran
Huang, Huaibo
Artificial Intelligence
While progress in GUI agents has been largely driven by industrial-scale training, ungrounded hallucinations often trigger cascading failures in real-world deployments.Unlike general VLM domains, the GUI agent field lacks a hallucination-focused suite for fine-grained diagnosis, reliable evaluation, and targeted mitigation.To bridge this gap, we introduce HalluClear, a comprehensive suite for hallucination mitigation in GUI agents as a complement to computation-intensive scaling. HalluClear comprises: (1) a GUI-specific hallucination taxonomy derived from empirical failure analysis; (2) a calibrated three-stage evaluation workflow which enhances VLM-as-a-judge reliability via expert-annotated benchmarking and ensemble credibility estimation; and (3) a mitigation scheme based on closed-loop structured reasoning, enabling lightweight continual post-training with cold-start initialization for both generalist and GUI-specialist agents. Experiments across representative agents and public benchmarks demonstrate that post-training on only 9K samples within our suite can significantly reduce hallucinations, thereby improving grounding and action fidelity, offering a compute-efficient pathway to robust GUI automation.
title HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2604.17284