Latent Space Disentanglement in Diffusion Transformers Enables Zero-shot Fine-grained Semantic Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Shuai, Zitao, Wu, Chenwei, Tang, Zhengxu, Song, Bowen, Shen, Liyue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Space Disentanglement in Diffusion Transformers Enables Precise Zero-shot Semantic Editing
by: Shuai, Zitao, et al.
Published: (2024)
by: Shuai, Zitao, et al.
Published: (2024)
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
by: Shuai, Zitao, et al.
Published: (2024)
by: Shuai, Zitao, et al.
Published: (2024)
Efficient In-Context Medical Segmentation with Meta-driven Visual Prompt Selection
by: Wu, Chenwei, et al.
Published: (2024)
by: Wu, Chenwei, et al.
Published: (2024)
SatDiffMoE: A Mixture of Estimation Method for Satellite Image Super-resolution with Latent Diffusion Models
by: Luo, Zhaoxu, et al.
Published: (2024)
by: Luo, Zhaoxu, et al.
Published: (2024)
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
by: Song, James, et al.
Published: (2026)
by: Song, James, et al.
Published: (2026)
Solving Inverse Problems with Latent Diffusion Models via Hard Data Consistency
by: Song, Bowen, et al.
Published: (2023)
by: Song, Bowen, et al.
Published: (2023)
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers
by: Dalva, Yusuf, et al.
Published: (2024)
by: Dalva, Yusuf, et al.
Published: (2024)
Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection
by: Zhu, Jiawen, et al.
Published: (2024)
by: Zhu, Jiawen, et al.
Published: (2024)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
by: Akdemir, Kiymet, et al.
Published: (2025)
by: Akdemir, Kiymet, et al.
Published: (2025)
DiffusionBlend: Learning 3D Image Prior through Position-aware Diffusion Score Blending for 3D Computed Tomography Reconstruction
by: Song, Bowen, et al.
Published: (2024)
by: Song, Bowen, et al.
Published: (2024)
LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
by: Zhang, Zhenghao, et al.
Published: (2025)
by: Zhang, Zhenghao, et al.
Published: (2025)
SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models
by: Tang, Zhengxu, et al.
Published: (2025)
by: Tang, Zhengxu, et al.
Published: (2025)
CoSIGN: Few-Step Guidance of ConSIstency Model to Solve General INverse Problems
by: Zhao, Jiankun, et al.
Published: (2024)
by: Zhao, Jiankun, et al.
Published: (2024)
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
AnimeAdapter: Fine-grained and Consistent Zero-shot Anime Character Generation
by: Han, Yixuan
Published: (2026)
by: Han, Yixuan
Published: (2026)
Learning Image Priors through Patch-based Diffusion Models for Solving Inverse Problems
by: Hu, Jason, et al.
Published: (2024)
by: Hu, Jason, et al.
Published: (2024)
Patch-Based Diffusion Models Beat Whole-Image Models for Mismatched Distribution Inverse Problems
by: Hu, Jason, et al.
Published: (2024)
by: Hu, Jason, et al.
Published: (2024)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
by: Shao, Dian, et al.
Published: (2026)
by: Shao, Dian, et al.
Published: (2026)
Latent Diffusion Inversion Requires Understanding the Latent Space
by: Rao, Mingxing, et al.
Published: (2025)
by: Rao, Mingxing, et al.
Published: (2025)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
by: Wang, Zhecan, et al.
Published: (2023)
by: Wang, Zhecan, et al.
Published: (2023)
DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing
by: Jia, Haozhe, et al.
Published: (2023)
by: Jia, Haozhe, et al.
Published: (2023)
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
by: Xu, Yu, et al.
Published: (2025)
by: Xu, Yu, et al.
Published: (2025)
MotionDiff: Training-free Zero-shot Interactive Motion Editing via Flow-assisted Multi-view Diffusion
by: Ma, Yikun, et al.
Published: (2025)
by: Ma, Yikun, et al.
Published: (2025)
Image-to-Image Translation with Disentangled Latent Vectors for Face Editing
by: Dalva, Yusuf, et al.
Published: (2023)
by: Dalva, Yusuf, et al.
Published: (2023)
FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
KPNDepth: Depth Estimation of Lane Images under Complex Rainy Environment
by: Shi, Zhengxu
Published: (2024)
by: Shi, Zhengxu
Published: (2024)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
Exploring Iterative Manifold Constraint for Zero-shot Image Editing
by: Li, Maomao, et al.
Published: (2025)
by: Li, Maomao, et al.
Published: (2025)
Boosting Latent Diffusion Models via Disentangled Representation Alignment
by: Page, John, et al.
Published: (2026)
by: Page, John, et al.
Published: (2026)
Part-aware Prompted Segment Anything Model for Adaptive Segmentation
by: Zhao, Chenhui, et al.
Published: (2024)
by: Zhao, Chenhui, et al.
Published: (2024)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion
by: Zeng, Yu, et al.
Published: (2024)
by: Zeng, Yu, et al.
Published: (2024)
Graph-guided Cross-composition Feature Disentanglement for Compositional Zero-shot Learning
by: Geng, Yuxia, et al.
Published: (2024)
by: Geng, Yuxia, et al.
Published: (2024)
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Similar Items
-
Latent Space Disentanglement in Diffusion Transformers Enables Precise Zero-shot Semantic Editing
by: Shuai, Zitao, et al.
Published: (2024) -
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
by: Shuai, Zitao, et al.
Published: (2024) -
Efficient In-Context Medical Segmentation with Meta-driven Visual Prompt Selection
by: Wu, Chenwei, et al.
Published: (2024) -
SatDiffMoE: A Mixture of Estimation Method for Satellite Image Super-resolution with Latent Diffusion Models
by: Luo, Zhaoxu, et al.
Published: (2024) -
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
by: Song, James, et al.
Published: (2026)