Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dong In, Doh, Hyungjun, Chi, Seunggeun, Duan, Runlin, Kim, Sangpil, Ramani, Karthik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
by: Doh, Hyungjun, et al.
Published: (2025)
by: Doh, Hyungjun, et al.
Published: (2025)
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
by: Doh, Hyungjun, et al.
Published: (2025)
by: Doh, Hyungjun, et al.
Published: (2025)
CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
by: Roh, Wonseok, et al.
Published: (2024)
by: Roh, Wonseok, et al.
Published: (2024)
M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
by: Chi, Seunggeun, et al.
Published: (2024)
by: Chi, Seunggeun, et al.
Published: (2024)
An HCI-Centric Survey and Taxonomy of Human-Generative-AI Interactions
by: Shi, Jingyu, et al.
Published: (2023)
by: Shi, Jingyu, et al.
Published: (2023)
InfoGCN++: Learning Representation by Predicting the Future for Online Human Skeleton-based Action Recognition
by: Chi, Seunggeun, et al.
Published: (2023)
by: Chi, Seunggeun, et al.
Published: (2023)
CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence
by: Shi, Jingyu, et al.
Published: (2025)
by: Shi, Jingyu, et al.
Published: (2025)
DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products
by: Duan, Runlin, et al.
Published: (2025)
by: Duan, Runlin, et al.
Published: (2025)
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data
by: Chi, Seunggeun, et al.
Published: (2024)
by: Chi, Seunggeun, et al.
Published: (2024)
Visualizing Causality in Mixed Reality for Manual Task Learning: An Exploratory Study
by: Jain, Rahul, et al.
Published: (2023)
by: Jain, Rahul, et al.
Published: (2023)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025)
by: Lan, Rui, et al.
Published: (2025)
Towards Training-Free Scene Text Editing
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
by: Duan, Runlin, et al.
Published: (2025)
by: Duan, Runlin, et al.
Published: (2025)
Investigating Creativity in Humans and Generative AI Through Circles Exercises
by: Duan, Runlin, et al.
Published: (2025)
by: Duan, Runlin, et al.
Published: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024)
by: Lee, Keon, et al.
Published: (2024)
SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI
by: Duan, Runlin, et al.
Published: (2025)
by: Duan, Runlin, et al.
Published: (2025)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion Models
by: Han, Hyeonggeun, et al.
Published: (2025)
by: Han, Hyeonggeun, et al.
Published: (2025)
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
by: Yuan, Honghui, et al.
Published: (2025)
by: Yuan, Honghui, et al.
Published: (2025)
$Δ$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2024)
by: Chen, Pengtao, et al.
Published: (2024)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
by: Lee, Dong In, et al.
Published: (2024)
by: Lee, Dong In, et al.
Published: (2024)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
by: Li, Guangzhao, et al.
Published: (2025)
by: Li, Guangzhao, et al.
Published: (2025)
Edit Fidelity Field: Semantics-Aware Region Isolation for Training-Free Scene Text Editing
by: Li, Guandong, et al.
Published: (2026)
by: Li, Guandong, et al.
Published: (2026)
ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers
by: Huang, Lianghua, et al.
Published: (2024)
by: Huang, Lianghua, et al.
Published: (2024)
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
by: Li, Hongxi, et al.
Published: (2026)
by: Li, Hongxi, et al.
Published: (2026)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
TextFlux: An OCR‐Free DiT Model for High‐Fidelity Multilingual Scene Text Synthesis
by: Yu Xie, et al.
Published: (2026)
by: Yu Xie, et al.
Published: (2026)
Symmetry-Aware GFlowNets
by: Kim, Hohyun, et al.
Published: (2025)
by: Kim, Hohyun, et al.
Published: (2025)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024)
by: Zeng, Weichao, et al.
Published: (2024)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
by: Duan, Changxu
Published: (2025)
by: Duan, Changxu
Published: (2025)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
by: Choi, Saemee, et al.
Published: (2025)
by: Choi, Saemee, et al.
Published: (2025)
BlurGuard: A Simple Approach for Robustifying Image Protection Against AI-Powered Editing
by: Kim, Jinsu, et al.
Published: (2025)
by: Kim, Jinsu, et al.
Published: (2025)
On The Application of Linear Attention in Multimodal Transformers
by: Gerami, Armin, et al.
Published: (2026)
by: Gerami, Armin, et al.
Published: (2026)
Similar Items
-
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
by: Doh, Hyungjun, et al.
Published: (2025) -
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
by: Doh, Hyungjun, et al.
Published: (2025) -
CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
by: Roh, Wonseok, et al.
Published: (2024) -
M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
by: Chi, Seunggeun, et al.
Published: (2024) -
An HCI-Centric Survey and Taxonomy of Human-Generative-AI Interactions
by: Shi, Jingyu, et al.
Published: (2023)