Unified Human-Scene Interaction via Prompted Chain-of-Contacts
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Zeqi, Wang, Tai, Wang, Jingbo, Cao, Jinkun, Zhang, Wenwei, Dai, Bo, Lin, Dahua, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
by: Chen, Jiahe, et al.
Published: (2024)
by: Chen, Jiahe, et al.
Published: (2024)
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
by: Wang, Zhenzhi, et al.
Published: (2023)
by: Wang, Zhenzhi, et al.
Published: (2023)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
Multi-Object Tracking by Hierarchical Visual Representations
by: Cao, Jinkun, et al.
Published: (2024)
by: Cao, Jinkun, et al.
Published: (2024)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
by: Pan, Liang, et al.
Published: (2025)
by: Pan, Liang, et al.
Published: (2025)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
EAG-PT: Emission-Aware Gaussians and Path Tracing for Diffuse Indoor Scene Reconstruction and Editing
by: Yang, Xijie, et al.
Published: (2026)
by: Yang, Xijie, et al.
Published: (2026)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
Grounded 3D-LLM with Referent Tokens
by: Chen, Yilun, et al.
Published: (2024)
by: Chen, Yilun, et al.
Published: (2024)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025)
by: Mi, Boyu, et al.
Published: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
by: Hao, Jinkun, et al.
Published: (2026)
by: Hao, Jinkun, et al.
Published: (2026)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023)
by: Huang, Haifeng, et al.
Published: (2023)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026)
by: Huang, Haifeng, et al.
Published: (2026)
Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes
by: Jiang, Lihan, et al.
Published: (2024)
by: Jiang, Lihan, et al.
Published: (2024)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
by: Chen, Xiao, et al.
Published: (2024)
by: Chen, Xiao, et al.
Published: (2024)
HaloGS: Loose Coupling of Compact Geometry and Gaussian Splats for 3D Scenes
by: Jiang, Changjian, et al.
Published: (2025)
by: Jiang, Changjian, et al.
Published: (2025)
RoomTex: Texturing Compositional Indoor Scenes via Iterative Inpainting
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
by: Hao, Jinkun, et al.
Published: (2025)
by: Hao, Jinkun, et al.
Published: (2025)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
by: Zhu, Ruijie, et al.
Published: (2025)
by: Zhu, Ruijie, et al.
Published: (2025)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System
by: Luo, Zhen, et al.
Published: (2026)
by: Luo, Zhen, et al.
Published: (2026)
Harmony4D: A Video Dataset for In-The-Wild Close Human Interactions
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images
by: Dai, Jinkun, et al.
Published: (2026)
by: Dai, Jinkun, et al.
Published: (2026)
Controllable Text-to-Motion Generation via Modular Body-Part Phase Control
by: Dai, Minyue, et al.
Published: (2026)
by: Dai, Minyue, et al.
Published: (2026)
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
by: Wang, Chengfeng, et al.
Published: (2025)
by: Wang, Chengfeng, et al.
Published: (2025)
ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics
by: Gao, Jiawei, et al.
Published: (2024)
by: Gao, Jiawei, et al.
Published: (2024)
AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
by: Jiang, Lihan, et al.
Published: (2025)
by: Jiang, Lihan, et al.
Published: (2025)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
Similar Items
-
MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
by: Chen, Jiahe, et al.
Published: (2024) -
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
by: Wang, Zhenzhi, et al.
Published: (2023) -
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025) -
Multi-Object Tracking by Hierarchical Visual Representations
by: Cao, Jinkun, et al.
Published: (2024) -
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)