MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
Fuente:
arXiv
Salvato in:
| Autori principali: | Lyu, Ruiyuan, Lin, Jingli, Wang, Tai, Yang, Shuai, Mao, Xiaohan, Chen, Yilun, Xu, Runsen, Huang, Haifeng, Zhu, Chenming, Lin, Dahua, Pang, Jiangmiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024)
di: Xu, Runsen, et al.
Pubblicazione: (2024)
Grounded 3D-LLM with Referent Tokens
di: Chen, Yilun, et al.
Pubblicazione: (2024)
di: Chen, Yilun, et al.
Pubblicazione: (2024)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
di: Lin, Jingli, et al.
Pubblicazione: (2025)
di: Lin, Jingli, et al.
Pubblicazione: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
di: Hu, Miao, et al.
Pubblicazione: (2025)
di: Hu, Miao, et al.
Pubblicazione: (2025)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
di: Zhong, Weipeng, et al.
Pubblicazione: (2025)
di: Zhong, Weipeng, et al.
Pubblicazione: (2025)
PointLLM: Empowering Large Language Models to Understand Point Clouds
di: Xu, Runsen, et al.
Pubblicazione: (2023)
di: Xu, Runsen, et al.
Pubblicazione: (2023)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
di: Wei, Meng, et al.
Pubblicazione: (2025)
di: Wei, Meng, et al.
Pubblicazione: (2025)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
di: Gao, Ning, et al.
Pubblicazione: (2025)
di: Gao, Ning, et al.
Pubblicazione: (2025)
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation
di: Wei, Meng, et al.
Pubblicazione: (2025)
di: Wei, Meng, et al.
Pubblicazione: (2025)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
di: Wei, Meng, et al.
Pubblicazione: (2024)
di: Wei, Meng, et al.
Pubblicazione: (2024)
Learning H-Infinity Locomotion Control
di: Long, Junfeng, et al.
Pubblicazione: (2024)
di: Long, Junfeng, et al.
Pubblicazione: (2024)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
di: Xu, Xiaoxu, et al.
Pubblicazione: (2026)
di: Xu, Xiaoxu, et al.
Pubblicazione: (2026)
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
di: Yang, Sizhe, et al.
Pubblicazione: (2026)
di: Yang, Sizhe, et al.
Pubblicazione: (2026)
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
di: Yang, Sizhe, et al.
Pubblicazione: (2026)
di: Yang, Sizhe, et al.
Pubblicazione: (2026)
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
di: Tian, Yang, et al.
Pubblicazione: (2024)
di: Tian, Yang, et al.
Pubblicazione: (2024)
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
di: Ben, Qingwei, et al.
Pubblicazione: (2025)
di: Ben, Qingwei, et al.
Pubblicazione: (2025)
Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation
di: Yang, Sizhe, et al.
Pubblicazione: (2025)
di: Yang, Sizhe, et al.
Pubblicazione: (2025)
LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
di: Peng, Jiaqi, et al.
Pubblicazione: (2025)
di: Peng, Jiaqi, et al.
Pubblicazione: (2025)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
di: Lin, Jingli, et al.
Pubblicazione: (2025)
di: Lin, Jingli, et al.
Pubblicazione: (2025)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
di: Chen, Xiao, et al.
Pubblicazione: (2025)
di: Chen, Xiao, et al.
Pubblicazione: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
di: Peng, Jiaqi, et al.
Pubblicazione: (2025)
di: Peng, Jiaqi, et al.
Pubblicazione: (2025)
NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
di: Cai, Wenzhe, et al.
Pubblicazione: (2025)
di: Cai, Wenzhe, et al.
Pubblicazione: (2025)
Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains
di: Ben, Qingwei, et al.
Pubblicazione: (2025)
di: Ben, Qingwei, et al.
Pubblicazione: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
di: Wen, Xin, et al.
Pubblicazione: (2025)
di: Wen, Xin, et al.
Pubblicazione: (2025)
FedRC: A Rapid-Converged Hierarchical Federated Learning Framework in Street Scene Semantic Understanding
di: Kou, Wei-Bin, et al.
Pubblicazione: (2024)
di: Kou, Wei-Bin, et al.
Pubblicazione: (2024)
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
di: Han, Xiaoshen, et al.
Pubblicazione: (2025)
di: Han, Xiaoshen, et al.
Pubblicazione: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
di: Huang, Haifeng, et al.
Pubblicazione: (2023)
di: Huang, Haifeng, et al.
Pubblicazione: (2023)
UniCon: A Unified System for Efficient Robot Learning Transfers
di: Lin, Yunfeng, et al.
Pubblicazione: (2026)
di: Lin, Yunfeng, et al.
Pubblicazione: (2026)
HELIOS: Hierarchical Exploration for Language-Grounded Interaction in Open Scenes
di: Ashton, Katrina, et al.
Pubblicazione: (2025)
di: Ashton, Katrina, et al.
Pubblicazione: (2025)
HGACNet: Hierarchical Graph Attention Network for Cross-Modal Point Cloud Completion
di: Zeng, Yadan, et al.
Pubblicazione: (2025)
di: Zeng, Yadan, et al.
Pubblicazione: (2025)
Language-Grounded Hierarchical Planning and Execution with Multi-Robot 3D Scene Graphs
di: Strader, Jared, et al.
Pubblicazione: (2025)
di: Strader, Jared, et al.
Pubblicazione: (2025)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
di: Chen, Xiao, et al.
Pubblicazione: (2024)
di: Chen, Xiao, et al.
Pubblicazione: (2024)
Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control
di: Peng, Quanquan, et al.
Pubblicazione: (2026)
di: Peng, Quanquan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024) -
Grounded 3D-LLM with Referent Tokens
di: Chen, Yilun, et al.
Pubblicazione: (2024) -
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
di: Lin, Jingli, et al.
Pubblicazione: (2025) -
ChangingGrounding: 3D Visual Grounding in Changing Scenes
di: Hu, Miao, et al.
Pubblicazione: (2025) -
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
di: Zhong, Weipeng, et al.
Pubblicazione: (2025)