HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908978063081472 |
|---|---|
| author | Jin, Chao Yang, Wenkui Sun, Hao Liao, Yuqi Jiang, Qianyi Zhou, Kai Cao, Jie He, Ran Huang, Huaibo |
| author_facet | Jin, Chao Yang, Wenkui Sun, Hao Liao, Yuqi Jiang, Qianyi Zhou, Kai Cao, Jie He, Ran Huang, Huaibo |
| contents | While progress in GUI agents has been largely driven by industrial-scale training, ungrounded hallucinations often trigger cascading failures in real-world deployments.Unlike general VLM domains, the GUI agent field lacks a hallucination-focused suite for fine-grained diagnosis, reliable evaluation, and targeted mitigation.To bridge this gap, we introduce HalluClear, a comprehensive suite for hallucination mitigation in GUI agents as a complement to computation-intensive scaling. HalluClear comprises: (1) a GUI-specific hallucination taxonomy derived from empirical failure analysis; (2) a calibrated three-stage evaluation workflow which enhances VLM-as-a-judge reliability via expert-annotated benchmarking and ensemble credibility estimation; and (3) a mitigation scheme based on closed-loop structured reasoning, enabling lightweight continual post-training with cold-start initialization for both generalist and GUI-specialist agents. Experiments across representative agents and public benchmarks demonstrate that post-training on only 9K samples within our suite can significantly reduce hallucinations, thereby improving grounding and action fidelity, offering a compute-efficient pathway to robust GUI automation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_17284 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents Jin, Chao Yang, Wenkui Sun, Hao Liao, Yuqi Jiang, Qianyi Zhou, Kai Cao, Jie He, Ran Huang, Huaibo Artificial Intelligence While progress in GUI agents has been largely driven by industrial-scale training, ungrounded hallucinations often trigger cascading failures in real-world deployments.Unlike general VLM domains, the GUI agent field lacks a hallucination-focused suite for fine-grained diagnosis, reliable evaluation, and targeted mitigation.To bridge this gap, we introduce HalluClear, a comprehensive suite for hallucination mitigation in GUI agents as a complement to computation-intensive scaling. HalluClear comprises: (1) a GUI-specific hallucination taxonomy derived from empirical failure analysis; (2) a calibrated three-stage evaluation workflow which enhances VLM-as-a-judge reliability via expert-annotated benchmarking and ensemble credibility estimation; and (3) a mitigation scheme based on closed-loop structured reasoning, enabling lightweight continual post-training with cold-start initialization for both generalist and GUI-specialist agents. Experiments across representative agents and public benchmarks demonstrate that post-training on only 9K samples within our suite can significantly reduce hallucinations, thereby improving grounding and action fidelity, offering a compute-efficient pathway to robust GUI automation. |
| title | HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2604.17284 |