Gespeichert in:
| Hauptverfasser: | Zhang, Yifei, Zhao, Hao, Li, Hongyang, Chen, Siheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.08770 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
von: Ren, Shunli, et al.
Veröffentlicht: (2023)
von: Ren, Shunli, et al.
Veröffentlicht: (2023)
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
von: Kang, Shuhao, et al.
Veröffentlicht: (2023)
von: Kang, Shuhao, et al.
Veröffentlicht: (2023)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping
von: Liang, William, et al.
Veröffentlicht: (2026)
von: Liang, William, et al.
Veröffentlicht: (2026)
MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
von: Tang, Chao, et al.
Veröffentlicht: (2025)
von: Tang, Chao, et al.
Veröffentlicht: (2025)
OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
von: Guo, Jun, et al.
Veröffentlicht: (2026)
von: Guo, Jun, et al.
Veröffentlicht: (2026)
Learning from Massive Human Videos for Universal Humanoid Pose Control
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
Leveraging Unknown Objects to Construct Labeled-Unlabeled Meta-Relationships for Zero-Shot Object Navigation
von: Zheng, Yanwei, et al.
Veröffentlicht: (2024)
von: Zheng, Yanwei, et al.
Veröffentlicht: (2024)
Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images
von: Adrian, David B., et al.
Veröffentlicht: (2024)
von: Adrian, David B., et al.
Veröffentlicht: (2024)
Into the Unknown: Towards using Generative Models for Sampling Priors of Environment Uncertainty for Planning in Configuration Spaces
von: Bhattacharjee, Subhransu S., et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Subhransu S., et al.
Veröffentlicht: (2025)
End-to-end Autonomous Driving: Challenges and Frontiers
von: Chen, Li, et al.
Veröffentlicht: (2023)
von: Chen, Li, et al.
Veröffentlicht: (2023)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
SCENES: Subpixel Correspondence Estimation With Epipolar Supervision
von: Kloepfer, Dominik A., et al.
Veröffentlicht: (2024)
von: Kloepfer, Dominik A., et al.
Veröffentlicht: (2024)
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
von: Chang, Yue, et al.
Veröffentlicht: (2026)
von: Chang, Yue, et al.
Veröffentlicht: (2026)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
von: Chi, Haozhuang, et al.
Veröffentlicht: (2026)
von: Chi, Haozhuang, et al.
Veröffentlicht: (2026)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
Visual SLAMMOT Considering Multiple Motion Models
von: Tian, Peilin, et al.
Veröffentlicht: (2024)
von: Tian, Peilin, et al.
Veröffentlicht: (2024)
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
von: Feng, Zhicheng, et al.
Veröffentlicht: (2025)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
von: Li, Heng, et al.
Veröffentlicht: (2024)
von: Li, Heng, et al.
Veröffentlicht: (2024)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning
von: Shin, Ukcheol, et al.
Veröffentlicht: (2025)
von: Shin, Ukcheol, et al.
Veröffentlicht: (2025)
Fast maneuver recovery from aerial observation: trajectory clustering and outliers rejection
von: de Moura, Nelson, et al.
Veröffentlicht: (2024)
von: de Moura, Nelson, et al.
Veröffentlicht: (2024)
FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions
von: Berenguel-Baeta, Bruno, et al.
Veröffentlicht: (2022)
von: Berenguel-Baeta, Bruno, et al.
Veröffentlicht: (2022)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
Chain of World: World Model Thinking in Latent Motion
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
von: Tang, Zecong, et al.
Veröffentlicht: (2026)
von: Tang, Zecong, et al.
Veröffentlicht: (2026)
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
von: Zhong, Fangwei, et al.
Veröffentlicht: (2024)
von: Zhong, Fangwei, et al.
Veröffentlicht: (2024)
PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
von: Jiang, Hanxiao, et al.
Veröffentlicht: (2025)
von: Jiang, Hanxiao, et al.
Veröffentlicht: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
von: Fujii, Ryo, et al.
Veröffentlicht: (2024) -
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
von: Lei, Zixing, et al.
Veröffentlicht: (2026) -
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
von: Ren, Shunli, et al.
Veröffentlicht: (2023) -
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
von: Kang, Shuhao, et al.
Veröffentlicht: (2023) -
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)