When Do Diffusion Models learn to Generate Multiple Objects?
Fuente:
arXiv
Saved in:
| Main Authors: | Jeong, Yujin, Uselis, Arnas, Laina, Iro, Oh, Seong Joon, Rohrbach, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025)
by: Jeong, Yujin, et al.
Published: (2025)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)
by: Kargi, Bora, et al.
Published: (2026)
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025)
by: Sonthalia, Ankit, et al.
Published: (2025)
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025)
by: Koishigarina, Darina, et al.
Published: (2025)
How can embedding models bind concepts?
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
by: Boduljak, Gabrijel, et al.
Published: (2025)
by: Boduljak, Gabrijel, et al.
Published: (2025)
Learning segmentation from point trajectories
by: Karazija, Laurynas, et al.
Published: (2025)
by: Karazija, Laurynas, et al.
Published: (2025)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
by: Kim, Nayeong, et al.
Published: (2025)
by: Kim, Nayeong, et al.
Published: (2025)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
by: Chen, Ziyuan, et al.
Published: (2026)
by: Chen, Ziyuan, et al.
Published: (2026)
TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
by: Lee, Sohyun, et al.
Published: (2025)
by: Lee, Sohyun, et al.
Published: (2025)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026)
by: Rodriguez, Hector G., et al.
Published: (2026)
Diffusion Models for Open-Vocabulary Segmentation
by: Karazija, Laurynas, et al.
Published: (2023)
by: Karazija, Laurynas, et al.
Published: (2023)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)
by: Zhang, Chenshuang, et al.
Published: (2025)
Universal Algorithm-Implicit Learning
by: Woerner, Stefano, et al.
Published: (2026)
by: Woerner, Stefano, et al.
Published: (2026)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
Conditional Diffusion Model for Longitudinal Medical Image Generation
by: Dao, Duy-Phuong, et al.
Published: (2024)
by: Dao, Duy-Phuong, et al.
Published: (2024)
DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting
by: Engstler, Paul, et al.
Published: (2024)
by: Engstler, Paul, et al.
Published: (2024)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
GuidNoise: Single-Pair Guided Diffusion for Generalized Noise Synthesis
by: Kim, Changjin, et al.
Published: (2025)
by: Kim, Changjin, et al.
Published: (2025)
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
by: Han, Yujin, et al.
Published: (2025)
by: Han, Yujin, et al.
Published: (2025)
Real-Time Person Image Synthesis Using a Flow Matching Model
by: Jeong, Jiwoo, et al.
Published: (2025)
by: Jeong, Jiwoo, et al.
Published: (2025)
Explorer: Robust Collection of Interactable GUI Elements
by: Chaimalas, Iason, et al.
Published: (2025)
by: Chaimalas, Iason, et al.
Published: (2025)
Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
by: Lee, Minseung, et al.
Published: (2024)
by: Lee, Minseung, et al.
Published: (2024)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
HOIGS: Human-Object Interaction Gaussian Splatting
by: Kim, Taewoo, et al.
Published: (2026)
by: Kim, Taewoo, et al.
Published: (2026)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment
by: Laina, Sebastián Barbas, et al.
Published: (2025)
by: Laina, Sebastián Barbas, et al.
Published: (2025)
Similar Items
-
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026) -
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026) -
On the rankability of visual embeddings
by: Sonthalia, Ankit, et al.
Published: (2025) -
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025)