E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Trung X., Kang, Zhang, Hong, Ji Woo, Zheng, Xuran, Yoo, Chang D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-view Masked Diffusion Transformers for Person Image Synthesis
by: Pham, Trung X., et al.
Published: (2024)
by: Pham, Trung X., et al.
Published: (2024)
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
by: Pham, Trung X., et al.
Published: (2026)
by: Pham, Trung X., et al.
Published: (2026)
MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation
by: Pham, Trung X., et al.
Published: (2024)
by: Pham, Trung X., et al.
Published: (2024)
ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On
by: Hong, Ji Woo, et al.
Published: (2025)
by: Hong, Ji Woo, et al.
Published: (2025)
DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving
by: Zheng, Xuran, et al.
Published: (2025)
by: Zheng, Xuran, et al.
Published: (2025)
Zero-Shot Dual-Path Integration Framework for Open-Vocabulary 3D Instance Segmentation
by: Ton, Tri, et al.
Published: (2024)
by: Ton, Tri, et al.
Published: (2024)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
DNI: Dilutional Noise Initialization for Diffusion Video Editing
by: Yoon, Sunjae, et al.
Published: (2024)
by: Yoon, Sunjae, et al.
Published: (2024)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
by: Wang, Yuancheng, et al.
Published: (2024)
by: Wang, Yuancheng, et al.
Published: (2024)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
by: Cai, Shengqu, et al.
Published: (2024)
by: Cai, Shengqu, et al.
Published: (2024)
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
by: Ton, Tri, et al.
Published: (2025)
by: Ton, Tri, et al.
Published: (2025)
Single Mask and Large Amplification Electrothermal Microgripper
by: Phuc Hong-Pham, et al.
Published: (2025)
by: Phuc Hong-Pham, et al.
Published: (2025)
MultiDreamer3D: Multi-concept 3D Customization with Concept-Aware Diffusion Guidance
by: Song, Wooseok, et al.
Published: (2025)
by: Song, Wooseok, et al.
Published: (2025)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
by: Le, Minh-Quan, et al.
Published: (2023)
by: Le, Minh-Quan, et al.
Published: (2023)
SemTra: A Semantic Skill Translator for Cross-Domain Zero-Shot Policy Adaptation
by: Shin, Sangwoo, et al.
Published: (2024)
by: Shin, Sangwoo, et al.
Published: (2024)
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
Infinite Mask Diffusion for Few-Step Distillation
by: Yoo, Jaehoon, et al.
Published: (2026)
by: Yoo, Jaehoon, et al.
Published: (2026)
Occlusion-robust Stylization for Drawing-based 3D Animation
by: Yoon, Sunjae, et al.
Published: (2025)
by: Yoon, Sunjae, et al.
Published: (2025)
FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing
by: Koo, Gwanhyeong, et al.
Published: (2024)
by: Koo, Gwanhyeong, et al.
Published: (2024)
FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields
by: Koo, Gwanhyeong, et al.
Published: (2025)
by: Koo, Gwanhyeong, et al.
Published: (2025)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
by: Vo, Dinh-Khoi, et al.
Published: (2026)
by: Vo, Dinh-Khoi, et al.
Published: (2026)
Handling Supervision Scarcity in Chest X-ray Classification: Long-Tailed and Zero-Shot Learning
by: Pham, Ha-Hieu, et al.
Published: (2026)
by: Pham, Ha-Hieu, et al.
Published: (2026)
Efficient Zero-Shot Inpainting with Decoupled Diffusion Guidance
by: Moufad, Badr, et al.
Published: (2025)
by: Moufad, Badr, et al.
Published: (2025)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
by: Lee, Junhyeok, et al.
Published: (2025)
by: Lee, Junhyeok, et al.
Published: (2025)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2025)
by: Jiang, Ziyue, et al.
Published: (2025)
DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
by: Gao, Sicheng, et al.
Published: (2025)
by: Gao, Sicheng, et al.
Published: (2025)
Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation
by: Chen, Changgu, et al.
Published: (2024)
by: Chen, Changgu, et al.
Published: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
by: Wu, Zhichao, et al.
Published: (2025)
by: Wu, Zhichao, et al.
Published: (2025)
LEARNING ACHIEVEMENT AND KNOWLEDGE TRANSFER: THE IMPACT FACTOR OF E-LEARNING SYSTEM AT BACH KHOA UNIVERSITY, VIETNAM
by: Quoc Trung Pham
Published: (2018)
by: Quoc Trung Pham
Published: (2018)
Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
by: Song, Chenxi, et al.
Published: (2025)
by: Song, Chenxi, et al.
Published: (2025)
VLCounter: Text-aware Visual Representation for Zero-Shot Object Counting
by: Kang, Seunggu, et al.
Published: (2023)
by: Kang, Seunggu, et al.
Published: (2023)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
DOZE: A Dataset for Open-Vocabulary Zero-Shot Object Navigation in Dynamic Environments
by: Ma, Ji, et al.
Published: (2024)
by: Ma, Ji, et al.
Published: (2024)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Similar Items
-
Cross-view Masked Diffusion Transformers for Person Image Synthesis
by: Pham, Trung X., et al.
Published: (2024) -
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
by: Pham, Trung X., et al.
Published: (2026) -
MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation
by: Pham, Trung X., et al.
Published: (2024) -
ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On
by: Hong, Ji Woo, et al.
Published: (2025) -
DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving
by: Zheng, Xuran, et al.
Published: (2025)