TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qihang, Wang, Yaxiong, Cheng, Lechao, Zhong, Zhun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
by: Shen, Jinjie, et al.
Published: (2025)
by: Shen, Jinjie, et al.
Published: (2025)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
by: Li, Haiyang, et al.
Published: (2025)
by: Li, Haiyang, et al.
Published: (2025)
Knowledge Swapping via Learning and Unlearning
by: Xing, Mingyu, et al.
Published: (2025)
by: Xing, Mingyu, et al.
Published: (2025)
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
by: Wang, Yaxiong, et al.
Published: (2025)
by: Wang, Yaxiong, et al.
Published: (2025)
Text-Driven Diffusion Model for Sign Language Production
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
by: Shen, Jinjie, et al.
Published: (2026)
by: Shen, Jinjie, et al.
Published: (2026)
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
by: Shen, Jinjie, et al.
Published: (2026)
by: Shen, Jinjie, et al.
Published: (2026)
AdaptiveDrag: Semantic-Driven Dragging on Diffusion-Based Image Editing
by: Chen, DuoSheng, et al.
Published: (2024)
by: Chen, DuoSheng, et al.
Published: (2024)
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
by: Pu, Xinyu, et al.
Published: (2025)
by: Pu, Xinyu, et al.
Published: (2025)
RotationDrag: Point-based Image Editing with Rotated Diffusion Features
by: Luo, Minxing, et al.
Published: (2024)
by: Luo, Minxing, et al.
Published: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
DragGaussian: Enabling Drag-style Manipulation on 3D Gaussian Representation
by: Shen, Sitian, et al.
Published: (2024)
by: Shen, Sitian, et al.
Published: (2024)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
by: Zhou, Junbao, et al.
Published: (2025)
by: Zhou, Junbao, et al.
Published: (2025)
StableDrag: Stable Dragging for Point-based Image Editing
by: Cui, Yutao, et al.
Published: (2024)
by: Cui, Yutao, et al.
Published: (2024)
DragNeXt: Rethinking Drag-Based Image Editing
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment
by: Liao, Sheng-Hao, et al.
Published: (2025)
by: Liao, Sheng-Hao, et al.
Published: (2025)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
FastDrag: Manipulate Anything in One Step
by: Zhao, Xuanjia, et al.
Published: (2024)
by: Zhao, Xuanjia, et al.
Published: (2024)
ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
by: Xia, Siwei, et al.
Published: (2025)
by: Xia, Siwei, et al.
Published: (2025)
Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance
by: Wang, Jiacheng, et al.
Published: (2024)
by: Wang, Jiacheng, et al.
Published: (2024)
LoopGaussian: Creating 3D Cinemagraph with Multi-view Images via Eulerian Motion Field
by: Li, Jiyang, et al.
Published: (2024)
by: Li, Jiyang, et al.
Published: (2024)
ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion Model
by: Zang, Qi, et al.
Published: (2024)
by: Zang, Qi, et al.
Published: (2024)
One Stone with Two Birds: A Null-Text-Null Frequency-Aware Diffusion Models for Text-Guided Image Inpainting
by: Liu, Haipeng, et al.
Published: (2025)
by: Liu, Haipeng, et al.
Published: (2025)
RealDrag: The First Dragging Benchmark with Real Target Image
by: Zafarani, Ahmad, et al.
Published: (2025)
by: Zafarani, Ahmad, et al.
Published: (2025)
InstantDrag: Improving Interactivity in Drag-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2024)
by: Shin, Joonghyuk, et al.
Published: (2024)
Prior-Constrained Association Learning for Fine-Grained Generalized Category Discovery
by: Wang, Menglin, et al.
Published: (2025)
by: Wang, Menglin, et al.
Published: (2025)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
by: Xia, Yingjie, et al.
Published: (2025)
by: Xia, Yingjie, et al.
Published: (2025)
Memory Consistency Guided Divide-and-Conquer Learning for Generalized Category Discovery
by: Tu, Yuanpeng, et al.
Published: (2024)
by: Tu, Yuanpeng, et al.
Published: (2024)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
Similar Items
-
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024) -
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
by: Shen, Jinjie, et al.
Published: (2025) -
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
by: Zhang, Yan, et al.
Published: (2025) -
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024) -
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)