When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Tao, Zhou, Chuhao, Zhao, Guangyu, Cao, Haozhi, Pu, Yewen, Yang, Jianfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911301829132288
author Wu, Tao
Zhou, Chuhao
Zhao, Guangyu
Cao, Haozhi
Pu, Yewen
Yang, Jianfei
author_facet Wu, Tao
Zhou, Chuhao
Zhao, Guangyu
Cao, Haozhi
Pu, Yewen
Yang, Jianfei
contents Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information to answer. In this work, we focus on a minimal requirement for EQA agents, abstention: knowing when to withhold an answer. From an initial study of 500 human queries, we find that 32.4% contain missing or underspecified context. Drawing on this initial study and cognitive theories of human communication errors, we derive five representative categories requiring abstention: actionability limitation, referential underspecification, preference dependence, information unavailability, and false presupposition. We augment OpenEQA by having annotators transform well-posed questions into ambiguous variants outlined by these categories. The resulting dataset, AbstainEQA, comprises 1,636 annotated abstention cases paired with 1,636 original OpenEQA instances for balanced evaluation. Evaluating on AbstainEQA, we find that even the best frontier model only attains 42.79% abstention recall, while humans achieve 91.17%. We also find that scaling, prompting, and reasoning only yield marginal gains, and that fine-tuned models overfit to textual cues. Together, these results position abstention as a fundamental prerequisite for reliable interaction in embodied settings and as a necessary basis for effective clarification.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
Wu, Tao
Zhou, Chuhao
Zhao, Guangyu
Cao, Haozhi
Pu, Yewen
Yang, Jianfei
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information to answer. In this work, we focus on a minimal requirement for EQA agents, abstention: knowing when to withhold an answer. From an initial study of 500 human queries, we find that 32.4% contain missing or underspecified context. Drawing on this initial study and cognitive theories of human communication errors, we derive five representative categories requiring abstention: actionability limitation, referential underspecification, preference dependence, information unavailability, and false presupposition. We augment OpenEQA by having annotators transform well-posed questions into ambiguous variants outlined by these categories. The resulting dataset, AbstainEQA, comprises 1,636 annotated abstention cases paired with 1,636 original OpenEQA instances for balanced evaluation. Evaluating on AbstainEQA, we find that even the best frontier model only attains 42.79% abstention recall, while humans achieve 91.17%. We also find that scaling, prompting, and reasoning only yield marginal gains, and that fine-tuned models overfit to textual cues. Together, these results position abstention as a fundamental prerequisite for reliable interaction in embodied settings and as a necessary basis for effective clarification.
title When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2512.04597