Dreamweaver: Learning Compositional World Models from Pixels
Fuente:
arXiv
Saved in:
| Main Authors: | Baek, Junyeob, Wu, Yi-Fu, Singh, Gautam, Ahn, Sungjin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
by: Yoon, Jaesik, et al.
Published: (2023)
by: Yoon, Jaesik, et al.
Published: (2023)
Learning to Theorize the World from Observation
by: Baek, Doojin, et al.
Published: (2026)
by: Baek, Doojin, et al.
Published: (2026)
Neural Language of Thought Models
by: Wu, Yi-Fu, et al.
Published: (2024)
by: Wu, Yi-Fu, et al.
Published: (2024)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
by: Kwon, Soonwoo, et al.
Published: (2025)
by: Kwon, Soonwoo, et al.
Published: (2025)
Beyond Pixel Histories: World Models with Persistent 3D State
by: Garcin, Samuel, et al.
Published: (2026)
by: Garcin, Samuel, et al.
Published: (2026)
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
by: Athalye, Ashay, et al.
Published: (2024)
by: Athalye, Ashay, et al.
Published: (2024)
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
by: Jung, Whie, et al.
Published: (2024)
by: Jung, Whie, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
by: Baek, Junyeob, et al.
Published: (2025)
by: Baek, Junyeob, et al.
Published: (2025)
DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
by: Tang, Jinzhou, et al.
Published: (2026)
by: Tang, Jinzhou, et al.
Published: (2026)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
From Pixels to Components: Eigenvector Masking for Visual Representation Learning
by: Bizeul, Alice, et al.
Published: (2025)
by: Bizeul, Alice, et al.
Published: (2025)
GCI-ViTAL: Gradual Confidence Improvement with Vision Transformers for Active Learning on Label Noise
by: Mots'oehli, Moseli, et al.
Published: (2024)
by: Mots'oehli, Moseli, et al.
Published: (2024)
Pixels to Play: A Foundation Model for 3D Gameplay
by: Yue, Yuguang, et al.
Published: (2025)
by: Yue, Yuguang, et al.
Published: (2025)
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
by: Jang, Wonbong, et al.
Published: (2026)
by: Jang, Wonbong, et al.
Published: (2026)
IPixMatch: Boost Semi-supervised Semantic Segmentation with Inter-Pixel Relation
by: Wu, Kebin, et al.
Published: (2024)
by: Wu, Kebin, et al.
Published: (2024)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
by: Mots'oehli, Moseli, et al.
Published: (2025)
by: Mots'oehli, Moseli, et al.
Published: (2025)
Federated Learning with Feedback Alignment
by: Baek, Incheol, et al.
Published: (2025)
by: Baek, Incheol, et al.
Published: (2025)
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
by: Sun, Xintian, et al.
Published: (2024)
by: Sun, Xintian, et al.
Published: (2024)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
by: Pal, Avik, et al.
Published: (2024)
by: Pal, Avik, et al.
Published: (2024)
SafeDreamer: Safe Reinforcement Learning with World Models
by: Huang, Weidong, et al.
Published: (2023)
by: Huang, Weidong, et al.
Published: (2023)
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
by: Fu, Yuwei, et al.
Published: (2024)
by: Fu, Yuwei, et al.
Published: (2024)
PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving
by: Wozniak, Maciej K., et al.
Published: (2025)
by: Wozniak, Maciej K., et al.
Published: (2025)
Parallelized Spatiotemporal Binding
by: Singh, Gautam, et al.
Published: (2024)
by: Singh, Gautam, et al.
Published: (2024)
When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery
by: Xie, Yiqun, et al.
Published: (2024)
by: Xie, Yiqun, et al.
Published: (2024)
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
CURLing the Dream: Contrastive Representations for World Modeling in Reinforcement Learning
by: Kich, Victor Augusto, et al.
Published: (2024)
by: Kich, Victor Augusto, et al.
Published: (2024)
Pixel-level Certified Explanations via Randomized Smoothing
by: Anani, Alaa, et al.
Published: (2025)
by: Anani, Alaa, et al.
Published: (2025)
Pixel-Wise Recognition for Holistic Surgical Scene Understanding
by: Ayobi, Nicolás, et al.
Published: (2024)
by: Ayobi, Nicolás, et al.
Published: (2024)
PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views
by: Fei, Xin, et al.
Published: (2024)
by: Fei, Xin, et al.
Published: (2024)
An attempt to generate new bridge types from latent space of PixelCNN
by: Zhang, Hongjun
Published: (2024)
by: Zhang, Hongjun
Published: (2024)
Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
by: Zheng, Guangze, et al.
Published: (2025)
by: Zheng, Guangze, et al.
Published: (2025)
Topology-Preserving Polygon Augmentation for Segmentation in Structured Visual Domains
by: Laudari, Sudip, et al.
Published: (2026)
by: Laudari, Sudip, et al.
Published: (2026)
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
by: Shang, Xinyi, et al.
Published: (2026)
by: Shang, Xinyi, et al.
Published: (2026)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
by: Ghaemi, Hafez, et al.
Published: (2025)
by: Ghaemi, Hafez, et al.
Published: (2025)
Similar Items
-
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
by: Yoon, Jaesik, et al.
Published: (2023) -
Learning to Theorize the World from Observation
by: Baek, Doojin, et al.
Published: (2026) -
Neural Language of Thought Models
by: Wu, Yi-Fu, et al.
Published: (2024) -
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
by: Kwon, Soonwoo, et al.
Published: (2025) -
Beyond Pixel Histories: World Models with Persistent 3D State
by: Garcin, Samuel, et al.
Published: (2026)