Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Lv, Zhengyao, Pan, Tianlin, Si, Chenyang, Chen, Zhaoxi, Zuo, Wangmeng, Liu, Ziwei, Wong, Kwan-Yee K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)
by: Yang, Ying, et al.
Published: (2026)
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
by: Lv, Zhengyao, et al.
Published: (2026)
by: Lv, Zhengyao, et al.
Published: (2026)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
DiverseAR: Boosting Diversity in Bitwise Autoregressive Image Generation
by: Yang, Ying, et al.
Published: (2025)
by: Yang, Ying, et al.
Published: (2025)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
by: Pan, Tianlin, et al.
Published: (2026)
by: Pan, Tianlin, et al.
Published: (2026)
FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
by: Dong, Zhe, et al.
Published: (2024)
by: Dong, Zhe, et al.
Published: (2024)
Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising
by: Li, Junyi, et al.
Published: (2024)
by: Li, Junyi, et al.
Published: (2024)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
by: Xie, Chaohao, et al.
Published: (2025)
by: Xie, Chaohao, et al.
Published: (2025)
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
CiPR: An Efficient Framework with Cross-instance Positive Relations for Generalized Category Discovery
by: Hao, Shaozhe, et al.
Published: (2023)
by: Hao, Shaozhe, et al.
Published: (2023)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
by: Cao, Yukang, et al.
Published: (2025)
by: Cao, Yukang, et al.
Published: (2025)
PhysX-3D: Physical-Grounded 3D Asset Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
by: Liu, Zixian, et al.
Published: (2026)
by: Liu, Zixian, et al.
Published: (2026)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
FreeInit: Bridging Initialization Gap in Video Diffusion Models
by: Wu, Tianxing, et al.
Published: (2023)
by: Wu, Tianxing, et al.
Published: (2023)
LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
Multi-modal Crowd Counting via a Broker Modality
by: Meng, Haoliang, et al.
Published: (2024)
by: Meng, Haoliang, et al.
Published: (2024)
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
by: Cao, Yukang, et al.
Published: (2026)
by: Cao, Yukang, et al.
Published: (2026)
ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model
by: Chen, Binghui, et al.
Published: (2024)
by: Chen, Binghui, et al.
Published: (2024)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
by: Yang, Feng, et al.
Published: (2025)
by: Yang, Feng, et al.
Published: (2025)
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
by: Xie, Haozhe, et al.
Published: (2023)
by: Xie, Haozhe, et al.
Published: (2023)
Generative Gaussian Splatting for Unbounded 3D City Generation
by: Xie, Haozhe, et al.
Published: (2024)
by: Xie, Haozhe, et al.
Published: (2024)
Compositional Generative Model of Unbounded 4D Cities
by: Xie, Haozhe, et al.
Published: (2025)
by: Xie, Haozhe, et al.
Published: (2025)
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
by: Wang, Ruihe, et al.
Published: (2024)
by: Wang, Ruihe, et al.
Published: (2024)
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
by: Yang, Shuya, et al.
Published: (2024)
by: Yang, Shuya, et al.
Published: (2024)
Multi-Modality Driven LoRA for Adverse Condition Depth Estimation
by: Yang, Guanglei, et al.
Published: (2024)
by: Yang, Guanglei, et al.
Published: (2024)
Similar Items
-
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
by: Lv, Zhengyao, et al.
Published: (2025) -
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
by: Lv, Zhengyao, et al.
Published: (2024) -
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
by: Lv, Zhengyao, et al.
Published: (2024) -
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025) -
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)