Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Gen, Zhao, Bo, Yang, Jianfei, Sevilla-Lara, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
by: Chen, Yixiang, et al.
Published: (2026)
by: Chen, Yixiang, et al.
Published: (2026)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
MaskPlanner: Learning-Based Object-Centric Motion Generation from 3D Point Clouds
by: Tiboni, Gabriele, et al.
Published: (2025)
by: Tiboni, Gabriele, et al.
Published: (2025)
Masked Gamma-SSL: Learning Uncertainty Estimation via Masked Image Modeling
by: Williams, David S. W., et al.
Published: (2024)
by: Williams, David S. W., et al.
Published: (2024)
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)
by: Wang, Xuzhi, et al.
Published: (2025)
GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
by: Katsumata, Kei, et al.
Published: (2025)
by: Katsumata, Kei, et al.
Published: (2025)
MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
by: Park, Soomin, et al.
Published: (2026)
by: Park, Soomin, et al.
Published: (2026)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
FLASH: Efficient Visuomotor Policy via Sparse Sampling
by: Bai, Jiaqi, et al.
Published: (2026)
by: Bai, Jiaqi, et al.
Published: (2026)
Masked Depth Modeling for Spatial Perception
by: Tan, Bin, et al.
Published: (2026)
by: Tan, Bin, et al.
Published: (2026)
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation
by: Deng, Yijie, et al.
Published: (2025)
by: Deng, Yijie, et al.
Published: (2025)
Unsupervised UAV 3D Trajectories Estimation with Sparse Point Clouds
by: Liang, Hanfang, et al.
Published: (2024)
by: Liang, Hanfang, et al.
Published: (2024)
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024)
by: Van Vo, Tuan, et al.
Published: (2024)
Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
by: Hu, Bin, et al.
Published: (2025)
by: Hu, Bin, et al.
Published: (2025)
Pre-Trained Masked Image Model for Mobile Robot Navigation
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
UMAD: Unsupervised Mask-Level Anomaly Detection for Autonomous Driving
by: Bogdoll, Daniel, et al.
Published: (2024)
by: Bogdoll, Daniel, et al.
Published: (2024)
Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction
by: Paidi, Santosh Kumar
Published: (2026)
by: Paidi, Santosh Kumar
Published: (2026)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Gradient-Driven 3D Segmentation and Affordance Transfer in Gaussian Splatting Using 2D Masks
by: Joseph, Joji, et al.
Published: (2024)
by: Joseph, Joji, et al.
Published: (2024)
InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling
by: Javed, Muhammad Gohar, et al.
Published: (2024)
by: Javed, Muhammad Gohar, et al.
Published: (2024)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Lift, Splat, Map: Lifting Foundation Masks for Label-Free Semantic Scene Completion
by: Zhang, Arthur, et al.
Published: (2024)
by: Zhang, Arthur, et al.
Published: (2024)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
by: Mascaro, Esteve Valls, et al.
Published: (2023)
by: Mascaro, Esteve Valls, et al.
Published: (2023)
MATT-GS: Masked Attention-based 3DGS for Robot Perception and Object Detection
by: Lee, Jee Won, et al.
Published: (2025)
by: Lee, Jee Won, et al.
Published: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
by: Jain, Yash, et al.
Published: (2023)
by: Jain, Yash, et al.
Published: (2023)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
by: Guan, Runwei, et al.
Published: (2026)
by: Guan, Runwei, et al.
Published: (2026)
DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
by: Tian, Jingyi, et al.
Published: (2025)
by: Tian, Jingyi, et al.
Published: (2025)
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024)
by: Zheng, Jinliang, et al.
Published: (2024)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
by: Yariv, Guy, et al.
Published: (2025)
by: Yariv, Guy, et al.
Published: (2025)
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
by: Zhang, Qifan, et al.
Published: (2026)
by: Zhang, Qifan, et al.
Published: (2026)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation
by: Sun, Zhengnan, et al.
Published: (2025)
by: Sun, Zhengnan, et al.
Published: (2025)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
by: Zhao, Hongxiang, et al.
Published: (2025)
by: Zhao, Hongxiang, et al.
Published: (2025)
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
by: Ma, Haoyu, et al.
Published: (2023)
by: Ma, Haoyu, et al.
Published: (2023)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Similar Items
-
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
by: Chen, Yixiang, et al.
Published: (2026) -
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024) -
MaskPlanner: Learning-Based Object-Centric Motion Generation from 3D Point Clouds
by: Tiboni, Gabriele, et al.
Published: (2025) -
Masked Gamma-SSL: Learning Uncertainty Estimation via Masked Image Modeling
by: Williams, David S. W., et al.
Published: (2024) -
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)