Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Joonghyuk, Hwang, Alchan, Kim, Yujin, Kim, Daneul, Park, Jaesik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstantDrag: Improving Interactivity in Drag-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2024)
by: Shin, Joonghyuk, et al.
Published: (2024)
Improving Editability in Image Generation with Layer-wise Memory
by: Kim, Daneul, et al.
Published: (2025)
by: Kim, Daneul, et al.
Published: (2025)
Pick-or-Mix: Dynamic Channel Sampling for ConvNets
by: Kumar, Ashish, et al.
Published: (2024)
by: Kumar, Ashish, et al.
Published: (2024)
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
by: Do, Seunguk, et al.
Published: (2026)
by: Do, Seunguk, et al.
Published: (2026)
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
by: Kwon, Mingi, et al.
Published: (2025)
by: Kwon, Mingi, et al.
Published: (2025)
Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
by: Kim, Daneul, et al.
Published: (2025)
by: Kim, Daneul, et al.
Published: (2025)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
Enhancing Creative Generation on Stable Diffusion-based Models
by: Han, Jiyeon, et al.
Published: (2025)
by: Han, Jiyeon, et al.
Published: (2025)
Leveraging Learned Image Prior for 3D Gaussian Compression
by: Shin, Seungjoo, et al.
Published: (2025)
by: Shin, Seungjoo, et al.
Published: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
When Model Knowledge meets Diffusion Model: Diffusion-assisted Data-free Image Synthesis with Alignment of Domain and Class
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Cross Resolution Encoding-Decoding For Detection Transformers
by: Kumar, Ashish, et al.
Published: (2024)
by: Kumar, Ashish, et al.
Published: (2024)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
by: Choi, Jaeyoung, et al.
Published: (2026)
by: Choi, Jaeyoung, et al.
Published: (2026)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
by: Kim, Sangmin, et al.
Published: (2026)
by: Kim, Sangmin, et al.
Published: (2026)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Metropolis-Hastings Sampling for 3D Gaussian Reconstruction
by: Kim, Hyunjin, et al.
Published: (2025)
by: Kim, Hyunjin, et al.
Published: (2025)
Distribution Matching Distillation without Fake Score Network
by: Kim, Youngjoong, et al.
Published: (2026)
by: Kim, Youngjoong, et al.
Published: (2026)
Targetless LiDAR-Camera Calibration with Neural Gaussian Splatting
by: Jung, Haebeom, et al.
Published: (2025)
by: Jung, Haebeom, et al.
Published: (2025)
Deep Cost Ray Fusion for Sparse Depth Video Completion
by: Kim, Jungeon, et al.
Published: (2024)
by: Kim, Jungeon, et al.
Published: (2024)
Locality-aware Gaussian Compression for Fast and High-quality Rendering
by: Shin, Seungjoo, et al.
Published: (2025)
by: Shin, Seungjoo, et al.
Published: (2025)
Improving Cone-Beam CT Image Quality with Knowledge Distillation-Enhanced Diffusion Model in Imbalanced Data Settings
by: Hwang, Joonil, et al.
Published: (2024)
by: Hwang, Joonil, et al.
Published: (2024)
MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
by: Kim, Kyungwon, et al.
Published: (2026)
by: Kim, Kyungwon, et al.
Published: (2026)
Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
by: Kim, Youngjoong, et al.
Published: (2026)
by: Kim, Youngjoong, et al.
Published: (2026)
Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
by: Chung, Minjae, et al.
Published: (2025)
by: Chung, Minjae, et al.
Published: (2025)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
DiffuseHigh: Training-free Progressive High-Resolution Image Synthesis through Structure Guidance
by: Kim, Younghyun, et al.
Published: (2024)
by: Kim, Younghyun, et al.
Published: (2024)
Extend3D: Town-Scale 3D Generation
by: Yoon, Seungwoo, et al.
Published: (2026)
by: Yoon, Seungwoo, et al.
Published: (2026)
Diffusion Model Compression for Image-to-Image Translation
by: Kim, Geonung, et al.
Published: (2024)
by: Kim, Geonung, et al.
Published: (2024)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
by: Choi, Saemee, et al.
Published: (2025)
by: Choi, Saemee, et al.
Published: (2025)
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
by: Kim, Jin Hyeon, et al.
Published: (2025)
by: Kim, Jin Hyeon, et al.
Published: (2025)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection
by: Lee, Chanhui, et al.
Published: (2026)
by: Lee, Chanhui, et al.
Published: (2026)
DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image Editing
by: Choi, June Suk, et al.
Published: (2024)
by: Choi, June Suk, et al.
Published: (2024)
Simulating Post-Neoadjuvant Chemotherapy Breast Cancer MRI via Diffusion Model with Prompt Tuning
by: Kim, Jonghun, et al.
Published: (2025)
by: Kim, Jonghun, et al.
Published: (2025)
Recovering Dynamic 3D Sketches from Videos
by: Lee, Jaeah, et al.
Published: (2025)
by: Lee, Jaeah, et al.
Published: (2025)
3Doodle: Compact Abstraction of Objects with 3D Strokes
by: Choi, Changwoon, et al.
Published: (2024)
by: Choi, Changwoon, et al.
Published: (2024)
Exploring Conditions for Diffusion models in Robotic Control
by: Shin, Heeseong, et al.
Published: (2025)
by: Shin, Heeseong, et al.
Published: (2025)
Designing Concise ConvNets with Columnar Stages
by: Kumar, Ashish, et al.
Published: (2024)
by: Kumar, Ashish, et al.
Published: (2024)
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
by: Kim, Hyeonyu, et al.
Published: (2025)
by: Kim, Hyeonyu, et al.
Published: (2025)
Similar Items
-
InstantDrag: Improving Interactivity in Drag-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2024) -
Improving Editability in Image Generation with Layer-wise Memory
by: Kim, Daneul, et al.
Published: (2025) -
Pick-or-Mix: Dynamic Channel Sampling for ConvNets
by: Kumar, Ashish, et al.
Published: (2024) -
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
by: Do, Seunguk, et al.
Published: (2026) -
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
by: Kwon, Mingi, et al.
Published: (2025)