Saved in:
| Main Author: | Paidi, Santosh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.15466 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
by: Li, Lei-lei, et al.
Published: (2025)
by: Li, Lei-lei, et al.
Published: (2025)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)
by: Wu, Meiqi, et al.
Published: (2026)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026)
by: Zhu, Shangwen, et al.
Published: (2026)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023)
by: Zadaianchuk, Andrii, et al.
Published: (2023)
GloPath: An Entity-Centric Foundation Model for Glomerular Lesion Assessment and Clinicopathological Insights
by: He, Qiming, et al.
Published: (2026)
by: He, Qiming, et al.
Published: (2026)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
by: Fang, Pengcheng, et al.
Published: (2025)
by: Fang, Pengcheng, et al.
Published: (2025)
WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes
by: Hu, Jichen, et al.
Published: (2026)
by: Hu, Jichen, et al.
Published: (2026)
ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Geometry-Aware Implicit Memory for Video World Models
by: Wei, Zhengxuan, et al.
Published: (2026)
by: Wei, Zhengxuan, et al.
Published: (2026)
LIVE: Long-horizon Interactive Video World Modeling
by: Huang, Junchao, et al.
Published: (2026)
by: Huang, Junchao, et al.
Published: (2026)
TextOCVP: Object-Centric Video Prediction with Language Guidance
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
SCP: Spatial Causal Prediction in Video
by: Zhao, Yanguang, et al.
Published: (2026)
by: Zhao, Yanguang, et al.
Published: (2026)
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
by: Jin, Siyoon, et al.
Published: (2025)
by: Jin, Siyoon, et al.
Published: (2025)
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026)
by: Xu, Xiaojie, et al.
Published: (2026)
Causal Physics Steering in Video World Models via Concept Activation Vectors
by: Alam, Nahid
Published: (2026)
by: Alam, Nahid
Published: (2026)
MFil-Mamba: Multi-Filter Scanning for Spatial Redundancy-Aware Visual State Space Models
by: Khadka, Puskal, et al.
Published: (2026)
by: Khadka, Puskal, et al.
Published: (2026)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
by: Vandersanden, Jente, et al.
Published: (2026)
by: Vandersanden, Jente, et al.
Published: (2026)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
RELIC: Interactive Video World Model with Long-Horizon Memory
by: Hong, Yicong, et al.
Published: (2025)
by: Hong, Yicong, et al.
Published: (2025)
PanSR: An Object-Centric Mask Transformer for Panoptic Segmentation
by: Žust, Lojze, et al.
Published: (2024)
by: Žust, Lojze, et al.
Published: (2024)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video
by: Fan, Junkai, et al.
Published: (2024)
by: Fan, Junkai, et al.
Published: (2024)
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
by: Xu, Longwei, et al.
Published: (2026)
by: Xu, Longwei, et al.
Published: (2026)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
by: Xi, Zeyu, et al.
Published: (2025)
by: Xi, Zeyu, et al.
Published: (2025)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
by: Wan, Zhang, et al.
Published: (2024)
by: Wan, Zhang, et al.
Published: (2024)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
by: Cheng, Ying, et al.
Published: (2025)
by: Cheng, Ying, et al.
Published: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
by: Ha, Hyeonjeong, et al.
Published: (2026)
by: Ha, Hyeonjeong, et al.
Published: (2026)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Data Collection-free Masked Video Modeling
by: Ishikawa, Yuchi, et al.
Published: (2024)
by: Ishikawa, Yuchi, et al.
Published: (2024)
Similar Items
-
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
by: Li, Gen, et al.
Published: (2025) -
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024) -
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025) -
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
by: Li, Lei-lei, et al.
Published: (2025) -
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)