Salvato in:
| Autori principali: | Zhang, Zhuofan, Zhu, Ziyu, Li, Junhao, Li, Pengxiang, Wang, Tianxu, Liu, Tengyu, Ma, Xiaojian, Chen, Yixin, Jia, Baoxiong, Huang, Siyuan, Li, Qing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2408.04034 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
di: Zhu, Ziyu, et al.
Pubblicazione: (2025)
di: Zhu, Ziyu, et al.
Pubblicazione: (2025)
Unifying 3D Vision-Language Understanding via Promptable Queries
di: Zhu, Ziyu, et al.
Pubblicazione: (2024)
di: Zhu, Ziyu, et al.
Pubblicazione: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
di: Linghu, Xiongkun, et al.
Pubblicazione: (2025)
di: Linghu, Xiongkun, et al.
Pubblicazione: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
di: Wang, Yan, et al.
Pubblicazione: (2025)
di: Wang, Yan, et al.
Pubblicazione: (2025)
Multi-modal Situated Reasoning in 3D Scenes
di: Linghu, Xiongkun, et al.
Pubblicazione: (2024)
di: Linghu, Xiongkun, et al.
Pubblicazione: (2024)
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
di: Yu, Huangyue, et al.
Pubblicazione: (2025)
di: Yu, Huangyue, et al.
Pubblicazione: (2025)
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
di: Wang, Zan, et al.
Pubblicazione: (2024)
di: Wang, Zan, et al.
Pubblicazione: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
di: Linghu, Xiongkun, et al.
Pubblicazione: (2026)
di: Linghu, Xiongkun, et al.
Pubblicazione: (2026)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
di: Yang, Yandan, et al.
Pubblicazione: (2024)
di: Yang, Yandan, et al.
Pubblicazione: (2024)
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
di: Yang, Yandan, et al.
Pubblicazione: (2025)
di: Yang, Yandan, et al.
Pubblicazione: (2025)
Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
An Embodied Generalist Agent in 3D World
di: Huang, Jiangyong, et al.
Pubblicazione: (2023)
di: Huang, Jiangyong, et al.
Pubblicazione: (2023)
Scaling Up Dynamic Human-Scene Interaction Modeling
di: Jiang, Nan, et al.
Pubblicazione: (2024)
di: Jiang, Nan, et al.
Pubblicazione: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
di: Liu, Yu, et al.
Pubblicazione: (2024)
di: Liu, Yu, et al.
Pubblicazione: (2024)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
di: Chen, Yixin, et al.
Pubblicazione: (2026)
di: Chen, Yixin, et al.
Pubblicazione: (2026)
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
di: Lu, Ruijie, et al.
Pubblicazione: (2024)
di: Lu, Ruijie, et al.
Pubblicazione: (2024)
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
di: Guo, Jun, et al.
Pubblicazione: (2024)
di: Guo, Jun, et al.
Pubblicazione: (2024)
Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
di: Wang, Zan, et al.
Pubblicazione: (2025)
di: Wang, Zan, et al.
Pubblicazione: (2025)
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
di: Li, Yuyang, et al.
Pubblicazione: (2025)
di: Li, Yuyang, et al.
Pubblicazione: (2025)
Dynamic Motion Blending for Versatile Motion Editing
di: Jiang, Nan, et al.
Pubblicazione: (2025)
di: Jiang, Nan, et al.
Pubblicazione: (2025)
GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill
di: Cui, Jieming, et al.
Pubblicazione: (2025)
di: Cui, Jieming, et al.
Pubblicazione: (2025)
Autonomous Character-Scene Interaction Synthesis from Text Instruction
di: Jiang, Nan, et al.
Pubblicazione: (2024)
di: Jiang, Nan, et al.
Pubblicazione: (2024)
Grasp Multiple Objects with One Hand
di: Li, Yuyang, et al.
Pubblicazione: (2023)
di: Li, Yuyang, et al.
Pubblicazione: (2023)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
di: Cui, Jieming, et al.
Pubblicazione: (2024)
di: Cui, Jieming, et al.
Pubblicazione: (2024)
3D Scene Change Modeling With Consistent Multi-View Aggregation
di: Zhou, Zirui, et al.
Pubblicazione: (2025)
di: Zhou, Zirui, et al.
Pubblicazione: (2025)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
di: Li, Kailin, et al.
Pubblicazione: (2025)
di: Li, Kailin, et al.
Pubblicazione: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
di: Li, Puhao, et al.
Pubblicazione: (2024)
di: Li, Puhao, et al.
Pubblicazione: (2024)
Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation
di: Li, Yuyang, et al.
Pubblicazione: (2025)
di: Li, Yuyang, et al.
Pubblicazione: (2025)
LessMimic: Long-Horizon Humanoid Interaction with Unified Distance Field Representations
di: Lin, Yutang, et al.
Pubblicazione: (2026)
di: Lin, Yutang, et al.
Pubblicazione: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
di: Zhu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Zhu, Xiaomeng, et al.
Pubblicazione: (2025)
PhyRecon: Physically Plausible Neural Scene Reconstruction
di: Ni, Junfeng, et al.
Pubblicazione: (2024)
di: Ni, Junfeng, et al.
Pubblicazione: (2024)
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
di: Gao, Zhi, et al.
Pubblicazione: (2024)
di: Gao, Zhi, et al.
Pubblicazione: (2024)
ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting
di: Liu, Yu, et al.
Pubblicazione: (2025)
di: Liu, Yu, et al.
Pubblicazione: (2025)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
di: Jiang, Wentao, et al.
Pubblicazione: (2025)
di: Jiang, Wentao, et al.
Pubblicazione: (2025)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
di: Ge, Luzhou, et al.
Pubblicazione: (2026)
di: Ge, Luzhou, et al.
Pubblicazione: (2026)
Documenti analoghi
-
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
di: Wang, Tianxu, et al.
Pubblicazione: (2025) -
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
di: Zhu, Ziyu, et al.
Pubblicazione: (2025) -
Unifying 3D Vision-Language Understanding via Promptable Queries
di: Zhu, Ziyu, et al.
Pubblicazione: (2024) -
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
di: Jia, Baoxiong, et al.
Pubblicazione: (2024) -
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
di: Linghu, Xiongkun, et al.
Pubblicazione: (2025)