Saved in:
| Main Authors: | Xu, Xinyu, Luo, Shengcheng, Yang, Yanchao, Li, Yong-Lu, Lu, Cewu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.14758 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisit Human-Scene Interaction via Space Occupancy
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023)
by: Wang, Ziyu, et al.
Published: (2023)
Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition
by: Luo, Shengcheng, et al.
Published: (2024)
by: Luo, Shengcheng, et al.
Published: (2024)
SemGrasp: Semantic Grasp Generation via Language Aligned Discretization
by: Li, Kailin, et al.
Published: (2024)
by: Li, Kailin, et al.
Published: (2024)
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
by: Ming, Yuhang, et al.
Published: (2026)
by: Ming, Yuhang, et al.
Published: (2026)
DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Take A Step Back: Rethinking the Two Stages in Visual Reasoning
by: Zhang, Mingyu, et al.
Published: (2024)
by: Zhang, Mingyu, et al.
Published: (2024)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
by: Lin, Bingqian, et al.
Published: (2025)
by: Lin, Bingqian, et al.
Published: (2025)
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
Multi-view Hand Reconstruction with a Point-Embedded Transformer
by: Yang, Lixin, et al.
Published: (2024)
by: Yang, Lixin, et al.
Published: (2024)
OmniCam: Unified Multimodal Video Generation via Camera Control
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
by: Wang, Xinkai, et al.
Published: (2026)
by: Wang, Xinkai, et al.
Published: (2026)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control
by: Zhang, Xiaohan, et al.
Published: (2024)
by: Zhang, Xiaohan, et al.
Published: (2024)
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
by: Peng, Weikun, et al.
Published: (2025)
by: Peng, Weikun, et al.
Published: (2025)
Nav-R1: Reasoning and Navigation in Embodied Scenes
by: Liu, Qingxiang, et al.
Published: (2025)
by: Liu, Qingxiang, et al.
Published: (2025)
OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
by: Zhan, Xinyu, et al.
Published: (2024)
by: Zhan, Xinyu, et al.
Published: (2024)
Controllable 3D Outdoor Scene Generation via Scene Graphs
by: Liu, Yuheng, et al.
Published: (2025)
by: Liu, Yuheng, et al.
Published: (2025)
QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2025)
by: Zuo, Sicheng, et al.
Published: (2025)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
by: Li, Yanjun, et al.
Published: (2025)
by: Li, Yanjun, et al.
Published: (2025)
Semantic Granularity Navigation in Image Editing
by: Lu, Liangsi, et al.
Published: (2026)
by: Lu, Liangsi, et al.
Published: (2026)
Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact Representation
by: Yu, Zhenjun, et al.
Published: (2024)
by: Yu, Zhenjun, et al.
Published: (2024)
Digital Gene: Learning about the Physical World through Analytic Concepts
by: Sun, Jianhua, et al.
Published: (2025)
by: Sun, Jianhua, et al.
Published: (2025)
Dense Policy: Bidirectional Autoregressive Learning of Actions
by: Su, Yue, et al.
Published: (2025)
by: Su, Yue, et al.
Published: (2025)
Informative Scene Graph Generation via Debiasing
by: Gao, Lianli, et al.
Published: (2023)
by: Gao, Lianli, et al.
Published: (2023)
ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving Augmentation
by: Bian, Siyuan, et al.
Published: (2024)
by: Bian, Siyuan, et al.
Published: (2024)
Kalib: Easy Hand-Eye Calibration with Reference Point Tracking
by: Tang, Tutian, et al.
Published: (2024)
by: Tang, Tutian, et al.
Published: (2024)
TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
NeuroMamba: Multi-Perspective Feature Interaction with Visual Mamba for Neuron Segmentation
by: Jiang, Liuyun, et al.
Published: (2026)
by: Jiang, Liuyun, et al.
Published: (2026)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
by: Wei, Jiude, et al.
Published: (2025)
by: Wei, Jiude, et al.
Published: (2025)
Distill Gold from Massive Ores: Bi-level Data Pruning towards Efficient Dataset Distillation
by: Xu, Yue, et al.
Published: (2023)
by: Xu, Yue, et al.
Published: (2023)
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
PALUM: Part-based Attention Learning for Unified Motion Retargeting
by: Liu, Siqi, et al.
Published: (2026)
by: Liu, Siqi, et al.
Published: (2026)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
by: Qiao, Yanyuan, et al.
Published: (2025)
by: Qiao, Yanyuan, et al.
Published: (2025)
Similar Items
-
Revisit Human-Scene Interaction via Space Occupancy
by: Liu, Xinpeng, et al.
Published: (2023) -
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023) -
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026) -
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023) -
Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition
by: Luo, Shengcheng, et al.
Published: (2024)