Can Diffusion Models Disentangle? A Theoretical Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Liming, Mirza, Muhammad Jehanzeb, Gong, Yishu, Gong, Yuan, Zhang, Jiaqi, Tracey, Brian H., Placek, Katerina, Vilela, Marco, Glass, James R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
by: Wang, Liming, et al.
Published: (2024)
by: Wang, Liming, et al.
Published: (2024)
Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
by: Jha, Saurav, et al.
Published: (2024)
by: Jha, Saurav, et al.
Published: (2024)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
by: Selch, Lukas, et al.
Published: (2025)
by: Selch, Lukas, et al.
Published: (2025)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
by: Hansen, Jacob, et al.
Published: (2025)
by: Hansen, Jacob, et al.
Published: (2025)
TTT-KD: Test-Time Training for 3D Semantic Segmentation through Knowledge Distillation from Foundation Models
by: Weijler, Lisa, et al.
Published: (2024)
by: Weijler, Lisa, et al.
Published: (2024)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
by: Araujo, Edson, et al.
Published: (2026)
by: Araujo, Edson, et al.
Published: (2026)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
by: Singh, Akshit, et al.
Published: (2025)
by: Singh, Akshit, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
Can We Talk Models Into Seeing the World Differently?
by: Gavrikov, Paul, et al.
Published: (2024)
by: Gavrikov, Paul, et al.
Published: (2024)
Into the Fog: Evaluating Robustness of Multiple Object Tracking
by: Kirillova, Nadezda, et al.
Published: (2024)
by: Kirillova, Nadezda, et al.
Published: (2024)
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
by: Jha, Saurav, et al.
Published: (2025)
by: Jha, Saurav, et al.
Published: (2025)
InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic Perspective
by: Zhang, Yuanhong, et al.
Published: (2025)
by: Zhang, Yuanhong, et al.
Published: (2025)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
by: Gavrikov, Paul, et al.
Published: (2025)
by: Gavrikov, Paul, et al.
Published: (2025)
Comparison Visual Instruction Tuning
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
by: Zou, Zihang, et al.
Published: (2026)
by: Zou, Zihang, et al.
Published: (2026)
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
by: Spoecklberger, Johannes, et al.
Published: (2025)
by: Spoecklberger, Johannes, et al.
Published: (2025)
Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?
by: Jia, Yuru, et al.
Published: (2025)
by: Jia, Yuru, et al.
Published: (2025)
Towards Multimodal In-Context Learning for Vision & Language Models
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Dynamic Eraser for Guided Concept Erasure in Diffusion Models
by: Gong, Qinghui
Published: (2026)
by: Gong, Qinghui
Published: (2026)
Continual Segmentation with Disentangled Objectness Learning and Class Recognition
by: Gong, Yizheng, et al.
Published: (2024)
by: Gong, Yizheng, et al.
Published: (2024)
Unleashing the Potential of Pre-Trained Diffusion Models for Generalizable Person Re-Identification
by: Li, Jiachen, et al.
Published: (2025)
by: Li, Jiachen, et al.
Published: (2025)
DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling
by: Xie, Xin, et al.
Published: (2024)
by: Xie, Xin, et al.
Published: (2024)
Robust Concept Erasure in Diffusion Models: A Theoretical Perspective on Security and Robustness
by: Fu, Zixuan, et al.
Published: (2025)
by: Fu, Zixuan, et al.
Published: (2025)
TauGenNet: Plasma-Driven Tau PET Image Synthesis via Text-Guided 3D Diffusion Models
by: Gong, Yuxin, et al.
Published: (2025)
by: Gong, Yuxin, et al.
Published: (2025)
ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement
by: Butt, Muhammad Atif, et al.
Published: (2024)
by: Butt, Muhammad Atif, et al.
Published: (2024)
Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
by: Long, Kaifang, et al.
Published: (2026)
by: Long, Kaifang, et al.
Published: (2026)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
by: Rouditchenko, Andrew, et al.
Published: (2024)
by: Rouditchenko, Andrew, et al.
Published: (2024)
DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
by: Huang, Enbo, et al.
Published: (2024)
by: Huang, Enbo, et al.
Published: (2024)
HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models
by: Xie, Xin, et al.
Published: (2026)
by: Xie, Xin, et al.
Published: (2026)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
by: Akdemir, Kiymet, et al.
Published: (2025)
by: Akdemir, Kiymet, et al.
Published: (2025)
Few-Shot Image Generation by Conditional Relaxing Diffusion Inversion
by: Cao, Yu, et al.
Published: (2024)
by: Cao, Yu, et al.
Published: (2024)
TextureDiffusion: Target Prompt Disentangled Editing for Various Texture Transfer
by: Su, Zihan, et al.
Published: (2024)
by: Su, Zihan, et al.
Published: (2024)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
by: Qi, Tianhao, et al.
Published: (2024)
by: Qi, Tianhao, et al.
Published: (2024)
Chain-of-Trajectories: Unlocking the Intrinsic Generative Optimality of Diffusion Models via Graph-Theoretic Planning
by: Chen, Ping, et al.
Published: (2026)
by: Chen, Ping, et al.
Published: (2026)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
by: Mehta, Videet, et al.
Published: (2026)
by: Mehta, Videet, et al.
Published: (2026)
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
by: Huang, Siteng, et al.
Published: (2023)
by: Huang, Siteng, et al.
Published: (2023)
Similar Items
-
Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
by: Wang, Liming, et al.
Published: (2024) -
Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
by: Jha, Saurav, et al.
Published: (2024) -
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
by: Selch, Lukas, et al.
Published: (2025) -
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
by: Hansen, Jacob, et al.
Published: (2025) -
TTT-KD: Test-Time Training for 3D Semantic Segmentation through Knowledge Distillation from Foundation Models
by: Weijler, Lisa, et al.
Published: (2024)