Diffusion for World Modeling: Visual Details Matter in Atari
Fuente:
arXiv
Saved in:
| Main Authors: | Alonso, Eloi, Jelley, Adam, Micheli, Vincent, Kanervisto, Anssi, Storkey, Amos, Pearce, Tim, Fleuret, François |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024)
by: Micheli, Vincent, et al.
Published: (2024)
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
by: Zhang, Weipu, et al.
Published: (2025)
by: Zhang, Weipu, et al.
Published: (2025)
Adversarial robustness of VAEs through the lens of local geometry
by: Khan, Asif, et al.
Published: (2022)
by: Khan, Asif, et al.
Published: (2022)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games
by: Schäfer, Lukas, et al.
Published: (2023)
by: Schäfer, Lukas, et al.
Published: (2023)
Restoring Real-World Images with an Internal Detail Enhancement Diffusion Model
by: Xiao, Peng, et al.
Published: (2025)
by: Xiao, Peng, et al.
Published: (2025)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
by: Saied, Youssef, et al.
Published: (2026)
by: Saied, Youssef, et al.
Published: (2026)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
by: Chen, Lifeng, et al.
Published: (2025)
by: Chen, Lifeng, et al.
Published: (2025)
Diffusion Models for Counterfactual Generation and Anomaly Detection in Brain Images
by: Fontanella, Alessandro, et al.
Published: (2023)
by: Fontanella, Alessandro, et al.
Published: (2023)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
by: Guo, Junliang, et al.
Published: (2025)
by: Guo, Junliang, et al.
Published: (2025)
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
by: Huang, Zitong, et al.
Published: (2026)
by: Huang, Zitong, et al.
Published: (2026)
Laminating Representation Autoencoders for Efficient Diffusion
by: Calvo-González, Ramón, et al.
Published: (2026)
by: Calvo-González, Ramón, et al.
Published: (2026)
CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
by: Zhou, Binjia, et al.
Published: (2025)
by: Zhou, Binjia, et al.
Published: (2025)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
by: Delfosse, Quentin, et al.
Published: (2023)
by: Delfosse, Quentin, et al.
Published: (2023)
From Structure to Detail: Hierarchical Distillation for Efficient Diffusion Model
by: Cheng, Hanbo, et al.
Published: (2025)
by: Cheng, Hanbo, et al.
Published: (2025)
DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing
by: Wang, Xiaolong, et al.
Published: (2024)
by: Wang, Xiaolong, et al.
Published: (2024)
DP-MDM: Detail-Preserving MR Reconstruction via Multiple Diffusion Models
by: Geng, Mengxiao, et al.
Published: (2024)
by: Geng, Mengxiao, et al.
Published: (2024)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
by: Cho, Jang Hyun, et al.
Published: (2025)
by: Cho, Jang Hyun, et al.
Published: (2025)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
by: Sun, Yujing, et al.
Published: (2025)
by: Sun, Yujing, et al.
Published: (2025)
Beyond Pixel Histories: World Models with Persistent 3D State
by: Garcin, Samuel, et al.
Published: (2026)
by: Garcin, Samuel, et al.
Published: (2026)
Few-Shot Learning with Class Imbalance
by: Ochal, Mateusz, et al.
Published: (2021)
by: Ochal, Mateusz, et al.
Published: (2021)
DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models
by: Tian, Mengyuan, et al.
Published: (2026)
by: Tian, Mengyuan, et al.
Published: (2026)
einspace: Searching for Neural Architectures from Fundamental Operations
by: Ericsson, Linus, et al.
Published: (2024)
by: Ericsson, Linus, et al.
Published: (2024)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
by: Zhang, Lijun, et al.
Published: (2026)
by: Zhang, Lijun, et al.
Published: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games
by: Krauss, Henrik, et al.
Published: (2026)
by: Krauss, Henrik, et al.
Published: (2026)
Mirage2Matter: A Physically Grounded Gaussian World Model from Video
by: Gao, Zhengqing, et al.
Published: (2026)
by: Gao, Zhengqing, et al.
Published: (2026)
Object Attribute Matters in Visual Question Answering
by: Li, Peize, et al.
Published: (2023)
by: Li, Peize, et al.
Published: (2023)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Language-Conditioned World Modeling for Visual Navigation
by: Dong, Yifei, et al.
Published: (2026)
by: Dong, Yifei, et al.
Published: (2026)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Guided Synthesis of Labeled Brain MRI Data Using Latent Diffusion Models for Segmentation of Enlarged Ventricles
by: Ruschke, Tim, et al.
Published: (2024)
by: Ruschke, Tim, et al.
Published: (2024)
Detail Reinforcement Diffusion Model: Augmentation Fine-Grained Visual Categorization in Few-Shot Conditions
by: Wu, Tianxu, et al.
Published: (2023)
by: Wu, Tianxu, et al.
Published: (2023)
Aligning Agents like Large Language Models
by: Jelley, Adam, et al.
Published: (2024)
by: Jelley, Adam, et al.
Published: (2024)
TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
by: Qin, Chunxia, et al.
Published: (2026)
by: Qin, Chunxia, et al.
Published: (2026)
Similar Items
-
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024) -
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
by: Zhang, Weipu, et al.
Published: (2025) -
Adversarial robustness of VAEs through the lens of local geometry
by: Khan, Asif, et al.
Published: (2022) -
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025) -
Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games
by: Schäfer, Lukas, et al.
Published: (2023)