Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Shreyam, Agrawal, P., Gupta, Priyam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatio-Temporal Graph Dual-Attention Network for Multi-Agent Prediction and Tracking
by: Li, Jiachen, et al.
Published: (2021)
by: Li, Jiachen, et al.
Published: (2021)
Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence
by: Gupta, Pranay, et al.
Published: (2025)
by: Gupta, Pranay, et al.
Published: (2025)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025)
by: Liu, Shaowei, et al.
Published: (2025)
Precise Mobile Manipulation of Small Everyday Objects
by: Gupta, Arjun, et al.
Published: (2025)
by: Gupta, Arjun, et al.
Published: (2025)
Sensor-Invariant Tactile Representation
by: Gupta, Harsh, et al.
Published: (2025)
by: Gupta, Harsh, et al.
Published: (2025)
Opening Articulated Structures in the Real World
by: Gupta, Arjun, et al.
Published: (2024)
by: Gupta, Arjun, et al.
Published: (2024)
Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
by: Lai, Delun, et al.
Published: (2025)
by: Lai, Delun, et al.
Published: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Exploring Transformer-Augmented LSTM for Temporal and Spatial Feature Learning in Trajectory Prediction
by: Raskoti, Chandra, et al.
Published: (2024)
by: Raskoti, Chandra, et al.
Published: (2024)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
Push Past Green: Learning to Look Behind Plant Foliage by Moving It
by: Zhang, Xiaoyu, et al.
Published: (2023)
by: Zhang, Xiaoyu, et al.
Published: (2023)
Adver-City: Open-Source Multi-Modal Dataset for Collaborative Perception Under Adverse Weather Conditions
by: Karvat, Mateus, et al.
Published: (2024)
by: Karvat, Mateus, et al.
Published: (2024)
Map Prediction and Generative Entropy for Multi-Agent Exploration
by: Spinos, Alexander, et al.
Published: (2025)
by: Spinos, Alexander, et al.
Published: (2025)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023)
by: Huang, Tao, et al.
Published: (2023)
Vision-based Multi-future Trajectory Prediction: A Survey
by: Huang, Renhao, et al.
Published: (2023)
by: Huang, Renhao, et al.
Published: (2023)
SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
by: Wu, Shengkai, et al.
Published: (2025)
by: Wu, Shengkai, et al.
Published: (2025)
M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
by: Udugama, U. V. B. L, et al.
Published: (2025)
by: Udugama, U. V. B. L, et al.
Published: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
Detecting and Mitigating System-Level Anomalies of Vision-Based Controllers
by: Gupta, Aryaman, et al.
Published: (2023)
by: Gupta, Aryaman, et al.
Published: (2023)
Recent Advances in Multi-Agent Human Trajectory Prediction: A Comprehensive Review
by: Finet, Céline, et al.
Published: (2025)
by: Finet, Céline, et al.
Published: (2025)
Tactile Modality Fusion for Vision-Language-Action Models
by: Morissette, Charlotte, et al.
Published: (2026)
by: Morissette, Charlotte, et al.
Published: (2026)
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
by: Long, Juncen, et al.
Published: (2025)
by: Long, Juncen, et al.
Published: (2025)
Learning Spatial Structure from Pre-Beamforming Per-Antenna Range-Doppler Radar Data via Visibility-Aware Cross-Modal Supervision
by: Sebastian, George, et al.
Published: (2026)
by: Sebastian, George, et al.
Published: (2026)
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024)
by: He, Haoran, et al.
Published: (2024)
Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning
by: Zhang, Zeqing, et al.
Published: (2024)
by: Zhang, Zeqing, et al.
Published: (2024)
Dynamic Aware: Adaptive Multi-Mode Out-of-Distribution Detection for Trajectory Prediction in Autonomous Vehicles
by: Guo, Tongfei, et al.
Published: (2025)
by: Guo, Tongfei, et al.
Published: (2025)
LiDAR-BIND-T: Improved and Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications
by: Balemans, Niels, et al.
Published: (2025)
by: Balemans, Niels, et al.
Published: (2025)
CorVS: Person Identification via Video Trajectory-Sensor Correspondence in a Real-World Warehouse
by: Kano, Kazuma, et al.
Published: (2025)
by: Kano, Kazuma, et al.
Published: (2025)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
by: Qi, Han, et al.
Published: (2025)
by: Qi, Han, et al.
Published: (2025)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling
by: Neumeier, Marion, et al.
Published: (2026)
by: Neumeier, Marion, et al.
Published: (2026)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
by: Khurana, Mehar, et al.
Published: (2024)
by: Khurana, Mehar, et al.
Published: (2024)
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
by: Howell, Owen Lewis, et al.
Published: (2025)
by: Howell, Owen Lewis, et al.
Published: (2025)
EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction via Polynomial Representations
by: Yao, Yue, et al.
Published: (2025)
by: Yao, Yue, et al.
Published: (2025)
ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers
by: Lygerakis, Fotios, et al.
Published: (2025)
by: Lygerakis, Fotios, et al.
Published: (2025)
Advancing Autonomous Driving: DepthSense with Radar and Spatial Attention
by: Hussain, Muhamamd Ishfaq, et al.
Published: (2021)
by: Hussain, Muhamamd Ishfaq, et al.
Published: (2021)
Spatio-Temporal Multi-Subgraph GCN for 3D Human Motion Prediction
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning
by: Zeng, Runhao, et al.
Published: (2024)
by: Zeng, Runhao, et al.
Published: (2024)
Similar Items
-
Spatio-Temporal Graph Dual-Attention Network for Multi-Agent Prediction and Tracking
by: Li, Jiachen, et al.
Published: (2021) -
Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence
by: Gupta, Pranay, et al.
Published: (2025) -
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024) -
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025) -
Precise Mobile Manipulation of Small Everyday Objects
by: Gupta, Arjun, et al.
Published: (2025)