Transformers and Slot Encoding for Sample Efficient Physical World Modelling
Fuente:
arXiv
Saved in:
| Main Authors: | Petri, Francesco, Asprino, Luigi, Gangemi, Aldo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Local Causal World Models with State Space Models and Attention
by: Petri, Francesco, et al.
Published: (2025)
by: Petri, Francesco, et al.
Published: (2025)
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
Object-Centric Learning with Slot Mixture Module
by: Kirilenko, Daniil, et al.
Published: (2023)
by: Kirilenko, Daniil, et al.
Published: (2023)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
by: Pramanik, Rishav, et al.
Published: (2024)
by: Pramanik, Rishav, et al.
Published: (2024)
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024)
by: Micheli, Vincent, et al.
Published: (2024)
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024)
by: Manasyan, Anna, et al.
Published: (2024)
VORTEX: Challenging CNNs at Texture Recognition by using Vision Transformers with Orderless and Randomized Token Encodings
by: Scabini, Leonardo, et al.
Published: (2025)
by: Scabini, Leonardo, et al.
Published: (2025)
PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series
by: Li, Haobo, et al.
Published: (2025)
by: Li, Haobo, et al.
Published: (2025)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
by: Lin, Juyi, et al.
Published: (2026)
by: Lin, Juyi, et al.
Published: (2026)
Slot Structured World Models
by: Collu, Jonathan, et al.
Published: (2024)
by: Collu, Jonathan, et al.
Published: (2024)
DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuning
by: Shi, Zhengxiang, et al.
Published: (2023)
by: Shi, Zhengxiang, et al.
Published: (2023)
Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free Applications
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
RadMamba: Efficient Human Activity Recognition through Radar-based Micro-Doppler-Oriented Mamba State-Space Model
by: Wu, Yizhuo, et al.
Published: (2025)
by: Wu, Yizhuo, et al.
Published: (2025)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production
by: Montero, David, et al.
Published: (2025)
by: Montero, David, et al.
Published: (2025)
Adversarial Examples in the Physical World: A Survey
by: Wang, Jiakai, et al.
Published: (2023)
by: Wang, Jiakai, et al.
Published: (2023)
Is Retain Set All You Need in Machine Unlearning? Restoring Performance of Unlearned Models with Out-Of-Distribution Images
by: Bonato, Jacopo, et al.
Published: (2024)
by: Bonato, Jacopo, et al.
Published: (2024)
Multi-Head Encoding for Extreme Label Classification
by: Liang, Daojun, et al.
Published: (2024)
by: Liang, Daojun, et al.
Published: (2024)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
by: Tang, Haotian, et al.
Published: (2024)
by: Tang, Haotian, et al.
Published: (2024)
Scaling Diffusion Transformers Efficiently via $μ$P
by: Zheng, Chenyu, et al.
Published: (2025)
by: Zheng, Chenyu, et al.
Published: (2025)
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
by: Sanghi, Aditya, et al.
Published: (2024)
by: Sanghi, Aditya, et al.
Published: (2024)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
by: Kwon, Soonwoo, et al.
Published: (2025)
by: Kwon, Soonwoo, et al.
Published: (2025)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026)
by: Lai, Yixuan, et al.
Published: (2026)
Identifiable Token Correspondence for World Models
by: Kim, Youngin, et al.
Published: (2026)
by: Kim, Youngin, et al.
Published: (2026)
World Modeling with Probabilistic Structure Integration
by: Kotar, Klemen, et al.
Published: (2025)
by: Kotar, Klemen, et al.
Published: (2025)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
Does a Neural Network Really Encode Symbolic Concepts?
by: Li, Mingjie, et al.
Published: (2023)
by: Li, Mingjie, et al.
Published: (2023)
Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
by: Xia, Zaishuo, et al.
Published: (2025)
by: Xia, Zaishuo, et al.
Published: (2025)
A Study on Context Length and Efficient Transformers for Biomedical Image Analysis
by: Hooper, Sarah M., et al.
Published: (2024)
by: Hooper, Sarah M., et al.
Published: (2024)
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
by: Olszewski, Jan, et al.
Published: (2023)
by: Olszewski, Jan, et al.
Published: (2023)
Similar Items
-
Learning Local Causal World Models with State Space Models and Attention
by: Petri, Francesco, et al.
Published: (2025) -
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025) -
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
by: Lao, Dong, et al.
Published: (2023) -
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025) -
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)