IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haochen, Zantout, Nader, Kachana, Pujith, Zhang, Ji, Wang, Wenshan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
by: Zhang, Haochen, et al.
Published: (2024)
by: Zhang, Haochen, et al.
Published: (2024)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)
by: Zhang, Jiaxi, et al.
Published: (2026)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
by: Bao, Muyi, et al.
Published: (2026)
by: Bao, Muyi, et al.
Published: (2026)
Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
by: Li, Samuel, et al.
Published: (2025)
by: Li, Samuel, et al.
Published: (2025)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation
by: Jiang, Haochen, et al.
Published: (2024)
by: Jiang, Haochen, et al.
Published: (2024)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
by: Guo, Xiyue, et al.
Published: (2025)
by: Guo, Xiyue, et al.
Published: (2025)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
by: Sun, Xuefei, et al.
Published: (2026)
by: Sun, Xuefei, et al.
Published: (2026)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
by: Guo, Heyu, et al.
Published: (2025)
by: Guo, Heyu, et al.
Published: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
by: Ye, Angen, et al.
Published: (2025)
by: Ye, Angen, et al.
Published: (2025)
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
by: Patel, Manthan, et al.
Published: (2025)
by: Patel, Manthan, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
by: Ge, Luzhou, et al.
Published: (2026)
by: Ge, Luzhou, et al.
Published: (2026)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026)
by: Hu, Xinggang, et al.
Published: (2026)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
by: Wu, Chenyang, et al.
Published: (2024)
by: Wu, Chenyang, et al.
Published: (2024)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
by: Rotondi, Dennis, et al.
Published: (2025)
by: Rotondi, Dennis, et al.
Published: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
by: Wen, Junjie, et al.
Published: (2025)
by: Wen, Junjie, et al.
Published: (2025)
OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
by: Kassab, Christina, et al.
Published: (2025)
by: Kassab, Christina, et al.
Published: (2025)
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
by: Wang, Xinkai, et al.
Published: (2026)
by: Wang, Xinkai, et al.
Published: (2026)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024)
by: Yang, Yuncong, et al.
Published: (2024)
MoD-SLAM: Monocular Dense Mapping for Unbounded 3D Scene Reconstruction
by: Zhou, Heng, et al.
Published: (2024)
by: Zhou, Heng, et al.
Published: (2024)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
Similar Items
-
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025) -
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
by: Zhang, Haochen, et al.
Published: (2024) -
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
by: Saxena, Pranav, et al.
Published: (2025) -
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025) -
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)