Latent Video Prediction Learns Better World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Alrasheed, Ali J, Parast, Aryan Yazdan, Azam, Basim, Bailey, James, Akhtar, Naveed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DDB: Diffusion Driven Balancing to Address Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2025)
by: Parast, Aryan Yazdan, et al.
Published: (2025)
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2026)
by: Parast, Aryan Yazdan, et al.
Published: (2026)
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
by: Parast, Aryan Yazdan, et al.
Published: (2025)
by: Parast, Aryan Yazdan, et al.
Published: (2025)
Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control
by: Azam, Basim, et al.
Published: (2025)
by: Azam, Basim, et al.
Published: (2025)
Suitability of KANs for Computer Vision: A preliminary investigation
by: Azam, Basim, et al.
Published: (2024)
by: Azam, Basim, et al.
Published: (2024)
Efficient Diffusion Models for Vision: A Survey
by: Ulhaq, Anwaar, et al.
Published: (2022)
by: Ulhaq, Anwaar, et al.
Published: (2022)
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Manipulating and Mitigating Generative Model Biases without Retraining
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Exploring Bias in over 100 Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2025)
by: Vice, Jordan, et al.
Published: (2025)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors
by: Yang, Peiyu, et al.
Published: (2026)
by: Yang, Peiyu, et al.
Published: (2026)
Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
by: Zhang, Licheng, et al.
Published: (2025)
by: Zhang, Licheng, et al.
Published: (2025)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes
by: Ibrahim, Muhammad, et al.
Published: (2025)
by: Ibrahim, Muhammad, et al.
Published: (2025)
DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
by: Zhang, Licheng, et al.
Published: (2025)
by: Zhang, Licheng, et al.
Published: (2025)
DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
by: Deng, Yueci, et al.
Published: (2026)
by: Deng, Yueci, et al.
Published: (2026)
Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion
by: Liang, Li, et al.
Published: (2025)
by: Liang, Li, et al.
Published: (2025)
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026)
by: Jiang, Le, et al.
Published: (2026)
Predictive but Not Plannable: RC-aux for Latent World Models
by: Li, Wenyuan, et al.
Published: (2026)
by: Li, Wenyuan, et al.
Published: (2026)
VRAG: Learning World Models for Interactive Video Generation
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
by: Lu, Tianyi, et al.
Published: (2023)
by: Lu, Tianyi, et al.
Published: (2023)
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)
by: Yang, Fuxiang, et al.
Published: (2026)
WorldModelBench: Judging Video Generation Models As World Models
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
Latent-Compressed Variational Autoencoder for Video Diffusion Models
by: Guan, Jiarui, et al.
Published: (2026)
by: Guan, Jiarui, et al.
Published: (2026)
Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs
by: Zou, Yanmei, et al.
Published: (2026)
by: Zou, Yanmei, et al.
Published: (2026)
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)
by: Joseph, Sonia, et al.
Published: (2026)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space
by: Zhu, Jian, et al.
Published: (2025)
by: Zhu, Jian, et al.
Published: (2025)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
V-SenseDrive: A Privacy-Preserving Road Video and In-Vehicle Sensor Fusion Framework for Road Safety & Driver Behaviour Modelling
by: Naveed, Muhammad, et al.
Published: (2025)
by: Naveed, Muhammad, et al.
Published: (2025)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement
by: Nzoyem, Roussel Desmond, et al.
Published: (2026)
by: Nzoyem, Roussel Desmond, et al.
Published: (2026)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
by: Darcet, Timothée, et al.
Published: (2025)
by: Darcet, Timothée, et al.
Published: (2025)
DiLA: Disentangled Latent Action World Models
by: Zhang, Tianqiu, et al.
Published: (2026)
by: Zhang, Tianqiu, et al.
Published: (2026)
A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery
by: Khan, Sarmad, et al.
Published: (2026)
by: Khan, Sarmad, et al.
Published: (2026)
Pre-Trained Video Generative Models as World Simulators
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
Similar Items
-
DDB: Diffusion Driven Balancing to Address Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2025) -
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2026) -
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
by: Parast, Aryan Yazdan, et al.
Published: (2025) -
Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control
by: Azam, Basim, et al.
Published: (2025) -
Suitability of KANs for Computer Vision: A preliminary investigation
by: Azam, Basim, et al.
Published: (2024)