Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Kaixuan, Liu, Yang, Chen, Weixing, Luo, Jingzhou, Chen, Ziliang, Pan, Ling, Li, Guanbin, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
by: Luo, Jingzhou, et al.
Published: (2025)
by: Luo, Jingzhou, et al.
Published: (2025)
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
by: Yin, Shicheng, et al.
Published: (2026)
by: Yin, Shicheng, et al.
Published: (2026)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
by: Song, Xinshuai, et al.
Published: (2024)
by: Song, Xinshuai, et al.
Published: (2024)
ODMixer: Fine-grained Spatial-temporal MLP for Metro Origin-Destination Prediction
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
by: Yin, Shicheng, et al.
Published: (2025)
by: Yin, Shicheng, et al.
Published: (2025)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Cross-Modal Causal Intervention for Medical Report Generation
by: Chen, Weixing, et al.
Published: (2023)
by: Chen, Weixing, et al.
Published: (2023)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
by: Frahm, Noah, et al.
Published: (2025)
by: Frahm, Noah, et al.
Published: (2025)
VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis
by: Yin, Shicheng, et al.
Published: (2024)
by: Yin, Shicheng, et al.
Published: (2024)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
by: Qian, Tianwen, et al.
Published: (2023)
by: Qian, Tianwen, et al.
Published: (2023)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
Visual Environment-Interactive Planning for Embodied Complex-Question Answering
by: Lan, Ning, et al.
Published: (2025)
by: Lan, Ning, et al.
Published: (2025)
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
by: Chen, Weixing, et al.
Published: (2026)
by: Chen, Weixing, et al.
Published: (2026)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
Map-based Modular Approach for Zero-shot Embodied Question Answering
by: Sakamoto, Koya, et al.
Published: (2024)
by: Sakamoto, Koya, et al.
Published: (2024)
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)
by: SR, Nikitha
Published: (2025)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
by: Chen, Xiuyuan, et al.
Published: (2023)
by: Chen, Xiuyuan, et al.
Published: (2023)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025)
by: Liu, Dayong, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
by: Zemskova, Tatiana, et al.
Published: (2026)
by: Zemskova, Tatiana, et al.
Published: (2026)
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning
by: Xu, Quanxing, et al.
Published: (2025)
by: Xu, Quanxing, et al.
Published: (2025)
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
by: Xie, Tianchi, et al.
Published: (2025)
by: Xie, Tianchi, et al.
Published: (2025)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
by: Wang, Zining, et al.
Published: (2025)
by: Wang, Zining, et al.
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Similar Items
-
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024) -
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
by: Luo, Jingzhou, et al.
Published: (2025) -
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025) -
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025) -
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
by: Yin, Shicheng, et al.
Published: (2026)