Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
Fuente:
arXiv
Saved in:
| Main Authors: | Pasca, Razvan-George, Gavryushin, Alexey, Hamza, Muhammad, Kuo, Yen-Ling, Mo, Kaichun, Van Gool, Luc, Hilliges, Otmar, Wang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
PALM: Predicting Actions through Language Models
by: Kim, Sanghwan, et al.
Published: (2023)
by: Kim, Sanghwan, et al.
Published: (2023)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction
by: Akbiyik, M. Eren, et al.
Published: (2023)
by: Akbiyik, M. Eren, et al.
Published: (2023)
GraspXL: Generating Grasping Motions for Diverse Objects at Scale
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
by: Ozdel, Suleyman, et al.
Published: (2024)
by: Ozdel, Suleyman, et al.
Published: (2024)
SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion
by: Ho, Hsuan-I, et al.
Published: (2023)
by: Ho, Hsuan-I, et al.
Published: (2023)
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
by: Lu, Feichi, et al.
Published: (2024)
by: Lu, Feichi, et al.
Published: (2024)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
Anticipating Future Object Compositions without Forgetting
by: Zahran, Youssef, et al.
Published: (2024)
by: Zahran, Youssef, et al.
Published: (2024)
WANDR: Intention-guided Human Motion Generation
by: Diomataris, Markos, et al.
Published: (2024)
by: Diomataris, Markos, et al.
Published: (2024)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling
by: Sadat, Seyedmorteza, et al.
Published: (2023)
by: Sadat, Seyedmorteza, et al.
Published: (2023)
Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
by: Mo, Qijie, et al.
Published: (2024)
by: Mo, Qijie, et al.
Published: (2024)
RHOBIN Challenge: Reconstruction of Human Object Interaction
by: Xie, Xianghui, et al.
Published: (2024)
by: Xie, Xianghui, et al.
Published: (2024)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023)
by: Motamed, Saman, et al.
Published: (2023)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
The Six Hug Commandments: Design and Evaluation of a Human-Sized Hugging Robot with Visual and Haptic Perception
by: Block, Alexis E., et al.
Published: (2021)
by: Block, Alexis E., et al.
Published: (2021)
In the Arms of a Robot: Designing Autonomous Hugging Robots with Intra-Hug Gestures
by: Block, Alexis E., et al.
Published: (2022)
by: Block, Alexis E., et al.
Published: (2022)
Human Hair Reconstruction with Strand-Aligned 3D Gaussians
by: Zakharov, Egor, et al.
Published: (2024)
by: Zakharov, Egor, et al.
Published: (2024)
StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
by: Savov, Nedko, et al.
Published: (2025)
by: Savov, Nedko, et al.
Published: (2025)
XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
by: Tan, Yuedong, et al.
Published: (2024)
by: Tan, Yuedong, et al.
Published: (2024)
Back to the Future: The Role of Past and Future Context Predictability in Incremental Language Production
by: Upadhye, Shiva, et al.
Published: (2025)
by: Upadhye, Shiva, et al.
Published: (2025)
ReLoo: Reconstructing Humans Dressed in Loose Garments from Monocular Video in the Wild
by: Guo, Chen, et al.
Published: (2024)
by: Guo, Chen, et al.
Published: (2024)
MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild
by: Jiang, Zeren, et al.
Published: (2024)
by: Jiang, Zeren, et al.
Published: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
Incremental Object Detection with Prompt-based Methods
by: Neuwirth-Trapp, Matthias, et al.
Published: (2025)
by: Neuwirth-Trapp, Matthias, et al.
Published: (2025)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
A Unified Approach for Text- and Image-guided 4D Scene Generation
by: Zheng, Yufeng, et al.
Published: (2023)
by: Zheng, Yufeng, et al.
Published: (2023)
RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
by: Neuwirth-Trapp, Matthias, et al.
Published: (2025)
by: Neuwirth-Trapp, Matthias, et al.
Published: (2025)
Regressor-Guided Generative Image Editing Balances User Emotions to Reduce Time Spent Online
by: Gebhardt, Christoph, et al.
Published: (2025)
by: Gebhardt, Christoph, et al.
Published: (2025)
ArtiGrasp: Physically Plausible Synthesis of Bi-Manual Dexterous Grasping and Articulation
by: Zhang, Hui, et al.
Published: (2023)
by: Zhang, Hui, et al.
Published: (2023)
Gaussian Garments: Reconstructing Simulation-Ready Clothing with Photorealistic Appearance from Multi-View Video
by: Rong, Boxiang, et al.
Published: (2024)
by: Rong, Boxiang, et al.
Published: (2024)
Contrastive Learning for Multi-Object Tracking with Transformers
by: De Plaen, Pierre-François, et al.
Published: (2023)
by: De Plaen, Pierre-François, et al.
Published: (2023)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
Similar Items
-
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025) -
PALM: Predicting Actions through Language Models
by: Kim, Sanghwan, et al.
Published: (2023) -
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024) -
Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction
by: Akbiyik, M. Eren, et al.
Published: (2023) -
GraspXL: Generating Grasping Motions for Diverse Objects at Scale
by: Zhang, Hui, et al.
Published: (2024)