Accurate and Efficient World Modeling with Masked Latent Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Burchi, Maxime, Timofte, Radu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025)
by: Jesani, Krunal, et al.
Published: (2025)
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026)
by: Adhikari, Santosh Premi, et al.
Published: (2026)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
Masked Image Modeling: A Survey
by: Hondru, Vlad, et al.
Published: (2024)
by: Hondru, Vlad, et al.
Published: (2024)
CBM: Curriculum by Masking
by: Jarca, Andrei, et al.
Published: (2024)
by: Jarca, Andrei, et al.
Published: (2024)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
by: Zhou, Jiaheng, et al.
Published: (2025)
by: Zhou, Jiaheng, et al.
Published: (2025)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Transformers and Slot Encoding for Sample Efficient Physical World Modelling
by: Petri, Francesco, et al.
Published: (2024)
by: Petri, Francesco, et al.
Published: (2024)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
by: Gado, Mohamed, et al.
Published: (2025)
by: Gado, Mohamed, et al.
Published: (2025)
CL-MAE: Curriculum-Learned Masked Autoencoders
by: Madan, Neelu, et al.
Published: (2023)
by: Madan, Neelu, et al.
Published: (2023)
Predictive but Not Plannable: RC-aux for Latent World Models
by: Li, Wenyuan, et al.
Published: (2026)
by: Li, Wenyuan, et al.
Published: (2026)
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
by: Zhang, Kevin, et al.
Published: (2022)
by: Zhang, Kevin, et al.
Published: (2022)
Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
by: Xia, Zaishuo, et al.
Published: (2025)
by: Xia, Zaishuo, et al.
Published: (2025)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
Probing the Latent World: Emergent Discrete Symbols and Physical Structure in Latent Representations
by: ming, Liu hung
Published: (2026)
by: ming, Liu hung
Published: (2026)
VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025)
by: Goyal, Sahil, et al.
Published: (2025)
Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
Learned Lightweight Smartphone ISP with Unpaired Data
by: Arhire, Andrei, et al.
Published: (2025)
by: Arhire, Andrei, et al.
Published: (2025)
From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs
by: Shrestha, Usha, et al.
Published: (2026)
by: Shrestha, Usha, et al.
Published: (2026)
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
by: Chiang, Jui-Che, et al.
Published: (2024)
by: Chiang, Jui-Che, et al.
Published: (2024)
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024)
by: Micheli, Vincent, et al.
Published: (2024)
MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
by: Saleem, Muhammad Usama, et al.
Published: (2024)
by: Saleem, Muhammad Usama, et al.
Published: (2024)
HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
by: Kumbong, Hermann, et al.
Published: (2025)
by: Kumbong, Hermann, et al.
Published: (2025)
Efficient Flow Matching using Latent Variables
by: Samaddar, Anirban, et al.
Published: (2025)
by: Samaddar, Anirban, et al.
Published: (2025)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
Leaf Angle Estimation using Mask R-CNN and LETR Vision Transformer
by: Margapuri, Venkat, et al.
Published: (2024)
by: Margapuri, Venkat, et al.
Published: (2024)
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
MMM: Generative Masked Motion Model
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
by: Ye, Hanrong, et al.
Published: (2023)
by: Ye, Hanrong, et al.
Published: (2023)
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
by: Kišš, Martin, et al.
Published: (2025)
by: Kišš, Martin, et al.
Published: (2025)
WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
by: Ali, Fayaz, et al.
Published: (2025)
by: Ali, Fayaz, et al.
Published: (2025)
Learning Visual Feature-Based World Models via Residual Latent Action
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Similar Items
-
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025) -
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024) -
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025) -
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026) -
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
by: Burchi, Maxime, et al.
Published: (2024)