Compositional Generative Modeling: A Single Model is Not All You Need
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yilun, Kaelbling, Leslie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
by: Athalye, Ashay, et al.
Published: (2024)
by: Athalye, Ashay, et al.
Published: (2024)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
DiMSam: Diffusion Models as Samplers for Task and Motion Planning under Partial Observability
by: Fang, Xiaolin, et al.
Published: (2023)
by: Fang, Xiaolin, et al.
Published: (2023)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
by: Shen, William, et al.
Published: (2023)
by: Shen, William, et al.
Published: (2023)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
Embodied Uncertainty-Aware Object Segmentation
by: Fang, Xiaolin, et al.
Published: (2024)
by: Fang, Xiaolin, et al.
Published: (2024)
Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
by: Bhatt, Neel P., et al.
Published: (2024)
by: Bhatt, Neel P., et al.
Published: (2024)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
ODIN: A Single Model for 2D and 3D Segmentation
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
by: Barrett, Samuel J., et al.
Published: (2025)
by: Barrett, Samuel J., et al.
Published: (2025)
Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards
by: Yu, Xuehui, et al.
Published: (2026)
by: Yu, Xuehui, et al.
Published: (2026)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
Keypoint Abstraction using Large Models for Object-Relative Imitation Learning
by: Fang, Xiaolin, et al.
Published: (2024)
by: Fang, Xiaolin, et al.
Published: (2024)
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models
by: Bharadwaj, Sagar, et al.
Published: (2026)
by: Bharadwaj, Sagar, et al.
Published: (2026)
One-Shot Manipulation Strategy Learning by Making Contact Analogies
by: Liu, Yuyao, et al.
Published: (2024)
by: Liu, Yuyao, et al.
Published: (2024)
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
by: Hu, Yafei, et al.
Published: (2023)
by: Hu, Yafei, et al.
Published: (2023)
Is Hyperbolic Space All You Need for Medical Anomaly Detection?
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
Is Intermediate Fusion All You Need for UAV-based Collaborative Perception?
by: Hao, Jiuwu, et al.
Published: (2025)
by: Hao, Jiuwu, et al.
Published: (2025)
PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics
by: Luan, Xueyu, et al.
Published: (2026)
by: Luan, Xueyu, et al.
Published: (2026)
Focus On What Matters: Separated Models For Visual-Based RL Generalization
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
SLEDGE: Synthesizing Driving Environments with Generative Models and Rule-Based Traffic
by: Chitta, Kashyap, et al.
Published: (2024)
by: Chitta, Kashyap, et al.
Published: (2024)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Is Retain Set All You Need in Machine Unlearning? Restoring Performance of Unlearned Models with Out-Of-Distribution Images
by: Bonato, Jacopo, et al.
Published: (2024)
by: Bonato, Jacopo, et al.
Published: (2024)
Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
by: Zhang, Haichao, et al.
Published: (2023)
by: Zhang, Haichao, et al.
Published: (2023)
Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC
by: Du, Yilun, et al.
Published: (2023)
by: Du, Yilun, et al.
Published: (2023)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024)
by: Xie, Zixun, et al.
Published: (2024)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
by: Qi, Han, et al.
Published: (2025)
by: Qi, Han, et al.
Published: (2025)
DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Generative Image as Action Models
by: Shridhar, Mohit, et al.
Published: (2024)
by: Shridhar, Mohit, et al.
Published: (2024)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
You Only Crash Once v2: Perceptually Consistent Strong Features for One-Stage Domain Adaptive Detection of Space Terrain
by: Chase Jr, Timothy, et al.
Published: (2025)
by: Chase Jr, Timothy, et al.
Published: (2025)
Real-World Robot Applications of Foundation Models: A Review
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
by: Qin, Yiran, et al.
Published: (2025)
by: Qin, Yiran, et al.
Published: (2025)
Fractional Diffusion Bridge Models
by: Nobis, Gabriel, et al.
Published: (2025)
by: Nobis, Gabriel, et al.
Published: (2025)
SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
by: Benarous, Elior, et al.
Published: (2025)
by: Benarous, Elior, et al.
Published: (2025)
Similar Items
-
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
by: Athalye, Ashay, et al.
Published: (2024) -
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024) -
DiMSam: Diffusion Models as Samplers for Task and Motion Planning under Partial Observability
by: Fang, Xiaolin, et al.
Published: (2023) -
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025) -
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
by: Shen, William, et al.
Published: (2023)