Saved in:
| Main Authors: | Zhang, Yifei, Zhao, Hao, Li, Hongyang, Chen, Siheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.08770 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
by: Ren, Shunli, et al.
Published: (2023)
by: Ren, Shunli, et al.
Published: (2023)
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
by: Kang, Shuhao, et al.
Published: (2023)
by: Kang, Shuhao, et al.
Published: (2023)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
by: Jiang, Haoran, et al.
Published: (2025)
by: Jiang, Haoran, et al.
Published: (2025)
Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping
by: Liang, William, et al.
Published: (2026)
by: Liang, William, et al.
Published: (2026)
MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
by: Tang, Chao, et al.
Published: (2025)
by: Tang, Chao, et al.
Published: (2025)
OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
by: Deng, Yinan, et al.
Published: (2024)
by: Deng, Yinan, et al.
Published: (2024)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)
by: Guo, Jun, et al.
Published: (2026)
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
Leveraging Unknown Objects to Construct Labeled-Unlabeled Meta-Relationships for Zero-Shot Object Navigation
by: Zheng, Yanwei, et al.
Published: (2024)
by: Zheng, Yanwei, et al.
Published: (2024)
Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images
by: Adrian, David B., et al.
Published: (2024)
by: Adrian, David B., et al.
Published: (2024)
Into the Unknown: Towards using Generative Models for Sampling Priors of Environment Uncertainty for Planning in Configuration Spaces
by: Bhattacharjee, Subhransu S., et al.
Published: (2025)
by: Bhattacharjee, Subhransu S., et al.
Published: (2025)
End-to-end Autonomous Driving: Challenges and Frontiers
by: Chen, Li, et al.
Published: (2023)
by: Chen, Li, et al.
Published: (2023)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
SCENES: Subpixel Correspondence Estimation With Epipolar Supervision
by: Kloepfer, Dominik A., et al.
Published: (2024)
by: Kloepfer, Dominik A., et al.
Published: (2024)
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
by: Chang, Yue, et al.
Published: (2026)
by: Chang, Yue, et al.
Published: (2026)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
by: Cui, Wenbo, et al.
Published: (2025)
by: Cui, Wenbo, et al.
Published: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
by: Chi, Haozhuang, et al.
Published: (2026)
by: Chi, Haozhuang, et al.
Published: (2026)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
Visual SLAMMOT Considering Multiple Motion Models
by: Tian, Peilin, et al.
Published: (2024)
by: Tian, Peilin, et al.
Published: (2024)
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
by: Feng, Zhicheng, et al.
Published: (2025)
by: Feng, Zhicheng, et al.
Published: (2025)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
by: Li, Heng, et al.
Published: (2024)
by: Li, Heng, et al.
Published: (2024)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning
by: Shin, Ukcheol, et al.
Published: (2025)
by: Shin, Ukcheol, et al.
Published: (2025)
Fast maneuver recovery from aerial observation: trajectory clustering and outliers rejection
by: de Moura, Nelson, et al.
Published: (2024)
by: de Moura, Nelson, et al.
Published: (2024)
FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions
by: Berenguel-Baeta, Bruno, et al.
Published: (2022)
by: Berenguel-Baeta, Bruno, et al.
Published: (2022)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
by: Hao, Haihong, et al.
Published: (2026)
by: Hao, Haihong, et al.
Published: (2026)
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)
by: Yang, Fuxiang, et al.
Published: (2026)
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
by: Lin, Zuyao, et al.
Published: (2026)
by: Lin, Zuyao, et al.
Published: (2026)
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
by: Yang, Jianing, et al.
Published: (2025)
by: Yang, Jianing, et al.
Published: (2025)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
by: Tang, Zecong, et al.
Published: (2026)
by: Tang, Zecong, et al.
Published: (2026)
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
by: Rotondi, Dennis, et al.
Published: (2025)
by: Rotondi, Dennis, et al.
Published: (2025)
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
by: Jiang, Hanxiao, et al.
Published: (2025)
by: Jiang, Hanxiao, et al.
Published: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Similar Items
-
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024) -
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
by: Lei, Zixing, et al.
Published: (2026) -
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
by: Ren, Shunli, et al.
Published: (2023) -
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
by: Kang, Shuhao, et al.
Published: (2023) -
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
by: Jiang, Haoran, et al.
Published: (2025)