X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jian, Peng, Qirong, Guo, Xu, Chen, Chen, Lu, Haonan, Yang, Zhenyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in non-English Text-to-Image Generation
by: Ma, Jian, et al.
Published: (2023)
by: Ma, Jian, et al.
Published: (2023)
GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models
by: Ma, Jian, et al.
Published: (2024)
by: Ma, Jian, et al.
Published: (2024)
Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
by: Ma, Jian, et al.
Published: (2023)
by: Ma, Jian, et al.
Published: (2023)
HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
Cross-Resolution Distribution Matching for Diffusion Distillation
by: Chen, Feiyang, et al.
Published: (2026)
by: Chen, Feiyang, et al.
Published: (2026)
Efficient Dataset Distillation via Minimax Diffusion
by: Gu, Jianyang, et al.
Published: (2023)
by: Gu, Jianyang, et al.
Published: (2023)
LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models
by: Zhang, Dingkun, et al.
Published: (2024)
by: Zhang, Dingkun, et al.
Published: (2024)
Simple and Fast Distillation of Diffusion Models
by: Zhou, Zhenyu, et al.
Published: (2024)
by: Zhou, Zhenyu, et al.
Published: (2024)
Understanding Generalization in Diffusion Distillation via Probability Flow Distance
by: Zhang, Huijie, et al.
Published: (2025)
by: Zhang, Huijie, et al.
Published: (2025)
FIAS: Feature Imbalance-Aware Medical Image Segmentation with Dynamic Fusion and Mixing Attention
by: Liu, Xiwei, et al.
Published: (2024)
by: Liu, Xiwei, et al.
Published: (2024)
SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation
by: Liu, Hongjian, et al.
Published: (2024)
by: Liu, Hongjian, et al.
Published: (2024)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Spiking Transformer:Introducing Accurate Addition-Only Spiking Self-Attention for Transformer
by: Guo, Yufei, et al.
Published: (2025)
by: Guo, Yufei, et al.
Published: (2025)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
by: Li, Handong, et al.
Published: (2025)
by: Li, Handong, et al.
Published: (2025)
Multimodal Dataset Distillation via Phased Teacher Models
by: Guo, Shengbin, et al.
Published: (2026)
by: Guo, Shengbin, et al.
Published: (2026)
CONCORD: Concept-Informed Diffusion for Dataset Distillation
by: Gu, Jianyang, et al.
Published: (2025)
by: Gu, Jianyang, et al.
Published: (2025)
Muon-Accelerated Attention Distillation for Real-Time Edge Synthesis via Optimized Latent Diffusion
by: Chen, Weiye, et al.
Published: (2025)
by: Chen, Weiye, et al.
Published: (2025)
Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
Efficient Diffusion Distillation via Embedding Loss
by: Ying, Jincheng, et al.
Published: (2026)
by: Ying, Jincheng, et al.
Published: (2026)
NaTex: Seamless Texture Generation as Latent Color Diffusion
by: Lai, Zeqiang, et al.
Published: (2025)
by: Lai, Zeqiang, et al.
Published: (2025)
AddSR: Accelerating Diffusion-based Blind Super-Resolution with Adversarial Diffusion Distillation
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
by: Ren, Xiaoming, et al.
Published: (2026)
by: Ren, Xiaoming, et al.
Published: (2026)
LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
Attention Distillation: A Unified Approach to Visual Characteristics Transfer
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
AMFD: Distillation via Adaptive Multimodal Fusion for Multispectral Pedestrian Detection
by: Chen, Zizhao, et al.
Published: (2024)
by: Chen, Zizhao, et al.
Published: (2024)
Fine-Grained Prototypes Distillation for Few-Shot Object Detection
by: Wang, Zichen, et al.
Published: (2024)
by: Wang, Zichen, et al.
Published: (2024)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
by: Shen, Tao, et al.
Published: (2025)
by: Shen, Tao, et al.
Published: (2025)
Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation
by: Yang, Chen
Published: (2026)
by: Yang, Chen
Published: (2026)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
by: Li, Haotang, et al.
Published: (2026)
by: Li, Haotang, et al.
Published: (2026)
Noise-Informed Diffusion-Generated Image Detection with Anomaly Attention
by: Guan, Weinan, et al.
Published: (2025)
by: Guan, Weinan, et al.
Published: (2025)
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
by: Bai, Haoran, et al.
Published: (2025)
by: Bai, Haoran, et al.
Published: (2025)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
by: Han, Chengxi, et al.
Published: (2024)
by: Han, Chengxi, et al.
Published: (2024)
SlimDiffSR: Toward Lightweight and Efficient Remote Sensing Image Super-Resolution via Diffusion Model Distillation
by: Wang, Ce, et al.
Published: (2026)
by: Wang, Ce, et al.
Published: (2026)
Similar Items
-
Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
by: Ma, Jian, et al.
Published: (2025) -
X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
by: Ma, Jian, et al.
Published: (2025) -
PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in non-English Text-to-Image Generation
by: Ma, Jian, et al.
Published: (2023) -
GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models
by: Ma, Jian, et al.
Published: (2024) -
Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
by: Ma, Jian, et al.
Published: (2023)