ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, Zekun, Dong, Runpei, Zhang, Shaochen, Geng, Haoran, Han, Chunrui, Ge, Zheng, Yi, Li, Ma, Kaisheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Positional Prompt Tuning for Efficient 3D Representation Learning
di: Zhang, Shaochen, et al.
Pubblicazione: (2024)
di: Zhang, Shaochen, et al.
Pubblicazione: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
di: Dong, Runpei, et al.
Pubblicazione: (2023)
di: Dong, Runpei, et al.
Pubblicazione: (2023)
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
di: Ye, Junliang, et al.
Pubblicazione: (2025)
di: Ye, Junliang, et al.
Pubblicazione: (2025)
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
di: Han, Chunrui, et al.
Pubblicazione: (2023)
di: Han, Chunrui, et al.
Pubblicazione: (2023)
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
di: Peng, Yuang, et al.
Pubblicazione: (2024)
di: Peng, Yuang, et al.
Pubblicazione: (2024)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
di: Xie, Yifan, et al.
Pubblicazione: (2025)
di: Xie, Yifan, et al.
Pubblicazione: (2025)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
di: Qi, Zekun, et al.
Pubblicazione: (2025)
di: Qi, Zekun, et al.
Pubblicazione: (2025)
Focus Anywhere for Fine-grained Multi-page Document Understanding
di: Liu, Chenglong, et al.
Pubblicazione: (2024)
di: Liu, Chenglong, et al.
Pubblicazione: (2024)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
di: Yu, Haoran, et al.
Pubblicazione: (2025)
di: Yu, Haoran, et al.
Pubblicazione: (2025)
S2DM: Sector-Shaped Diffusion Models for Video Generation
di: Lang, Haoran, et al.
Pubblicazione: (2024)
di: Lang, Haoran, et al.
Pubblicazione: (2024)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
di: Wu, Yuqi, et al.
Pubblicazione: (2024)
di: Wu, Yuqi, et al.
Pubblicazione: (2024)
Small Language Model Meets with Reinforced Vision Vocabulary
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
di: Chen, Jinyue, et al.
Pubblicazione: (2024)
di: Chen, Jinyue, et al.
Pubblicazione: (2024)
DAOcc: 3D Object Detection Assisted Multi-Sensor Fusion for 3D Occupancy Prediction
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
di: Dong, Runpei, et al.
Pubblicazione: (2026)
di: Dong, Runpei, et al.
Pubblicazione: (2026)
Accelerating Diffusion Models with One-to-Many Knowledge Distillation
di: Zhang, Linfeng, et al.
Pubblicazione: (2024)
di: Zhang, Linfeng, et al.
Pubblicazione: (2024)
Interact3D: Compositional 3D Generation of Interactive Objects
di: Shan, Hui, et al.
Pubblicazione: (2026)
di: Shan, Hui, et al.
Pubblicazione: (2026)
VSFormer: Mining Correlations in Flexible View Set for Multi-view 3D Shape Understanding
di: Sun, Hongyu, et al.
Pubblicazione: (2024)
di: Sun, Hongyu, et al.
Pubblicazione: (2024)
Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
di: Geng, Zichen, et al.
Pubblicazione: (2025)
di: Geng, Zichen, et al.
Pubblicazione: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
di: Tang, Haoran, et al.
Pubblicazione: (2025)
di: Tang, Haoran, et al.
Pubblicazione: (2025)
Perception in Reflection
di: Wei, Yana, et al.
Pubblicazione: (2025)
di: Wei, Yana, et al.
Pubblicazione: (2025)
ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling
di: Zhang, Shuyuan, et al.
Pubblicazione: (2025)
di: Zhang, Shuyuan, et al.
Pubblicazione: (2025)
HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation
di: Qu, Wentian, et al.
Pubblicazione: (2025)
di: Qu, Wentian, et al.
Pubblicazione: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
di: Yang, Panqi, et al.
Pubblicazione: (2025)
di: Yang, Panqi, et al.
Pubblicazione: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
di: Li, Zechuan, et al.
Pubblicazione: (2025)
di: Li, Zechuan, et al.
Pubblicazione: (2025)
PhysPart: Physically Plausible Part Completion for Interactable Objects
di: Luo, Rundong, et al.
Pubblicazione: (2024)
di: Luo, Rundong, et al.
Pubblicazione: (2024)
Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions
di: Dong, Ze, et al.
Pubblicazione: (2026)
di: Dong, Ze, et al.
Pubblicazione: (2026)
UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes
di: Geng, Zichen, et al.
Pubblicazione: (2025)
di: Geng, Zichen, et al.
Pubblicazione: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
di: Zheng, Ying, et al.
Pubblicazione: (2024)
di: Zheng, Ying, et al.
Pubblicazione: (2024)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
di: Qian, Zekun, et al.
Pubblicazione: (2026)
di: Qian, Zekun, et al.
Pubblicazione: (2026)
EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems
di: Ma, Xiao, et al.
Pubblicazione: (2025)
di: Ma, Xiao, et al.
Pubblicazione: (2025)
Geometric Point Attention Transformer for 3D Shape Reassembly
di: Li, Jiahan, et al.
Pubblicazione: (2024)
di: Li, Jiahan, et al.
Pubblicazione: (2024)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection
di: Jia, Mingda, et al.
Pubblicazione: (2024)
di: Jia, Mingda, et al.
Pubblicazione: (2024)
ContextHOI: Spatial Context Learning for Human-Object Interaction Detection
di: Jia, Mingda, et al.
Pubblicazione: (2024)
di: Jia, Mingda, et al.
Pubblicazione: (2024)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
di: Ma, Wufei, et al.
Pubblicazione: (2024)
di: Ma, Wufei, et al.
Pubblicazione: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
di: Wang, Yifei, et al.
Pubblicazione: (2025)
di: Wang, Yifei, et al.
Pubblicazione: (2025)
Boosting 3D Object Generation through PBR Materials
di: Wang, Yitong, et al.
Pubblicazione: (2024)
di: Wang, Yitong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Positional Prompt Tuning for Efficient 3D Representation Learning
di: Zhang, Shaochen, et al.
Pubblicazione: (2024) -
DreamLLM: Synergistic Multimodal Comprehension and Creation
di: Dong, Runpei, et al.
Pubblicazione: (2023) -
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
di: Ye, Junliang, et al.
Pubblicazione: (2025) -
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
di: Han, Chunrui, et al.
Pubblicazione: (2023) -
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
di: Peng, Yuang, et al.
Pubblicazione: (2024)