AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Huynh, Cuong, Popov, Maxim, Gridusov, Denis, Kolyubin, Sergey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction
di: Gridusov, Denis, et al.
Pubblicazione: (2026)
di: Gridusov, Denis, et al.
Pubblicazione: (2026)
GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization
di: Sidorov, Gennady, et al.
Pubblicazione: (2024)
di: Sidorov, Gennady, et al.
Pubblicazione: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024)
di: Xu, Runsen, et al.
Pubblicazione: (2024)
OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes
di: Kurkova, Regina, et al.
Pubblicazione: (2026)
di: Kurkova, Regina, et al.
Pubblicazione: (2026)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
di: Sun, Xuefei, et al.
Pubblicazione: (2026)
di: Sun, Xuefei, et al.
Pubblicazione: (2026)
Zero-Shot 3D Visual Grounding from Vision-Language Models
di: Li, Rong, et al.
Pubblicazione: (2025)
di: Li, Rong, et al.
Pubblicazione: (2025)
OSMa-Bench: Evaluating Open Semantic Mapping Under Varying Lighting Conditions
di: Popov, Maxim, et al.
Pubblicazione: (2025)
di: Popov, Maxim, et al.
Pubblicazione: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
di: Li, Rong, et al.
Pubblicazione: (2024)
di: Li, Rong, et al.
Pubblicazione: (2024)
Detecting the Anomalies in LiDAR Pointcloud
di: Zhang, Chiyu, et al.
Pubblicazione: (2023)
di: Zhang, Chiyu, et al.
Pubblicazione: (2023)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
di: Zantout, Nader, et al.
Pubblicazione: (2025)
di: Zantout, Nader, et al.
Pubblicazione: (2025)
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
di: Ci, Hai, et al.
Pubblicazione: (2025)
di: Ci, Hai, et al.
Pubblicazione: (2025)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
di: Kapelyukh, Ivan, et al.
Pubblicazione: (2023)
di: Kapelyukh, Ivan, et al.
Pubblicazione: (2023)
NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
di: Qin, Yiran, et al.
Pubblicazione: (2025)
di: Qin, Yiran, et al.
Pubblicazione: (2025)
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping
di: Iwase, Shun, et al.
Pubblicazione: (2025)
di: Iwase, Shun, et al.
Pubblicazione: (2025)
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
di: Zhang, Zixin, et al.
Pubblicazione: (2025)
di: Zhang, Zixin, et al.
Pubblicazione: (2025)
RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments
di: Nasser, Zaid, et al.
Pubblicazione: (2026)
di: Nasser, Zaid, et al.
Pubblicazione: (2026)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
di: Liu, Xuyang, et al.
Pubblicazione: (2023)
di: Liu, Xuyang, et al.
Pubblicazione: (2023)
Zero-Shot UAV Navigation in Forests via Relightable 3D Gaussian Splatting
di: Lv, Zinan, et al.
Pubblicazione: (2026)
di: Lv, Zinan, et al.
Pubblicazione: (2026)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
di: Zhang, Jiaxi, et al.
Pubblicazione: (2026)
di: Zhang, Jiaxi, et al.
Pubblicazione: (2026)
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
di: Newman, Benjamin A., et al.
Pubblicazione: (2024)
di: Newman, Benjamin A., et al.
Pubblicazione: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
di: Zhu, He, et al.
Pubblicazione: (2025)
di: Zhu, He, et al.
Pubblicazione: (2025)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
di: Tian, Tongxuan, et al.
Pubblicazione: (2025)
di: Tian, Tongxuan, et al.
Pubblicazione: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
di: Long, Zijun, et al.
Pubblicazione: (2023)
di: Long, Zijun, et al.
Pubblicazione: (2023)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
di: Chen, Kehan, et al.
Pubblicazione: (2024)
di: Chen, Kehan, et al.
Pubblicazione: (2024)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
di: Liu, Katherine, et al.
Pubblicazione: (2025)
di: Liu, Katherine, et al.
Pubblicazione: (2025)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
di: Saxena, Pranav, et al.
Pubblicazione: (2025)
di: Saxena, Pranav, et al.
Pubblicazione: (2025)
DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
di: Lu, Yiren, et al.
Pubblicazione: (2026)
di: Lu, Yiren, et al.
Pubblicazione: (2026)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
di: Huang, Xun, et al.
Pubblicazione: (2025)
di: Huang, Xun, et al.
Pubblicazione: (2025)
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
di: Yu, Ming-Ming, et al.
Pubblicazione: (2025)
di: Yu, Ming-Ming, et al.
Pubblicazione: (2025)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
di: Lian, Shiwei, et al.
Pubblicazione: (2024)
di: Lian, Shiwei, et al.
Pubblicazione: (2024)
DriveVA: Video Action Models are Zero-Shot Drivers
di: Liu, Mengmeng, et al.
Pubblicazione: (2026)
di: Liu, Mengmeng, et al.
Pubblicazione: (2026)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
di: Wang, Yixuan, et al.
Pubblicazione: (2023)
di: Wang, Yixuan, et al.
Pubblicazione: (2023)
ZeroSCD: Zero-Shot Street Scene Change Detection
di: Kannan, Shyam Sundar, et al.
Pubblicazione: (2024)
di: Kannan, Shyam Sundar, et al.
Pubblicazione: (2024)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
di: Patel, Naman, et al.
Pubblicazione: (2025)
di: Patel, Naman, et al.
Pubblicazione: (2025)
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
di: Jung, Seongmin, et al.
Pubblicazione: (2025)
di: Jung, Seongmin, et al.
Pubblicazione: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
di: Chen, Hanzhi, et al.
Pubblicazione: (2025)
di: Chen, Hanzhi, et al.
Pubblicazione: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
di: Guan, Tianrui, et al.
Pubblicazione: (2024)
di: Guan, Tianrui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction
di: Gridusov, Denis, et al.
Pubblicazione: (2026) -
GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization
di: Sidorov, Gennady, et al.
Pubblicazione: (2024) -
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024) -
OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes
di: Kurkova, Regina, et al.
Pubblicazione: (2026) -
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
di: Sun, Xuefei, et al.
Pubblicazione: (2026)