Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xinzhuo, Juvekar, Adheesh, Zhang, Jiaxun, Liu, Xingyou, Wahed, Muntasir, Nguyen, Kiet A., Shen, Yifan, Yu, Tianjiao, Lourentzou, Ismini
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918463514083328
author Li, Xinzhuo
Juvekar, Adheesh
Zhang, Jiaxun
Liu, Xingyou
Wahed, Muntasir
Nguyen, Kiet A.
Shen, Yifan
Yu, Tianjiao
Lourentzou, Ismini
author_facet Li, Xinzhuo
Juvekar, Adheesh
Zhang, Jiaxun
Liu, Xingyou
Wahed, Muntasir
Nguyen, Kiet A.
Shen, Yifan
Yu, Tianjiao
Lourentzou, Ismini
contents Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding hallucinations, producing masks for incorrect objects or for objects that are entirely absent. Existing evaluations rely almost entirely on text- or label-based perturbations, which check only whether the predicted mask matches the queried label. Such evaluations overlook the spatial footprint and severity of hallucination and therefore fail to reveal vision-driven hallucinations, which are more challenging and more prevalent. To address this gap, we formalize the task of Counterfactual Segmentation Reasoning (CSR), where a model must segment the referenced object in the factual image and abstain in its counterfactual counterpart. To support this task, we curate HalluSegBench, the first large-scale benchmark to diagnose referring and reasoning expression segmentation hallucinations using controlled visual counterfactuals, alongside new evaluation metrics that measure hallucination severity and disentangle vision- and language-driven failure modes. We further introduce RobustSeg, a segmentation VLM trained with counterfactual fine-tuning (CFT) to learn when to segment and when to abstain. Experimental results confirm RobustSeg reduces hallucinations by 30%, while improving segmentation performance on FP-RefCOCO(+/g).
format Preprint
id arxiv_https___arxiv_org_abs_2506_21546
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
Li, Xinzhuo
Juvekar, Adheesh
Zhang, Jiaxun
Liu, Xingyou
Wahed, Muntasir
Nguyen, Kiet A.
Shen, Yifan
Yu, Tianjiao
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding hallucinations, producing masks for incorrect objects or for objects that are entirely absent. Existing evaluations rely almost entirely on text- or label-based perturbations, which check only whether the predicted mask matches the queried label. Such evaluations overlook the spatial footprint and severity of hallucination and therefore fail to reveal vision-driven hallucinations, which are more challenging and more prevalent. To address this gap, we formalize the task of Counterfactual Segmentation Reasoning (CSR), where a model must segment the referenced object in the factual image and abstain in its counterfactual counterpart. To support this task, we curate HalluSegBench, the first large-scale benchmark to diagnose referring and reasoning expression segmentation hallucinations using controlled visual counterfactuals, alongside new evaluation metrics that measure hallucination severity and disentangle vision- and language-driven failure modes. We further introduce RobustSeg, a segmentation VLM trained with counterfactual fine-tuning (CFT) to learn when to segment and when to abstain. Experimental results confirm RobustSeg reduces hallucinations by 30%, while improving segmentation performance on FP-RefCOCO(+/g).
title Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.21546