A training-free framework for high-fidelity appearance transfer via diffusion transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Shengrong, Wang, Ye, Wu, Song, Ma, Rui, Wang, Qian, Wang, Lanjun, Yi, Zili
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912985482526720
author Gu, Shengrong
Wang, Ye
Wu, Song
Ma, Rui
Wang, Qian
Wang, Lanjun
Yi, Zili
author_facet Gu, Shengrong
Wang, Ye
Wu, Song
Ma, Rui
Wang, Qian
Wang, Lanjun
Yi, Zili
contents Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first training-free framework specifically designed to tame DiTs for high-fidelity appearance transfer. Our core is a synergistic system that disentangles structure and appearance. We leverage high-fidelity inversion to establish a rich content prior for the source image, capturing its lighting and micro-textures. A novel attention-sharing mechanism then dynamically fuses purified appearance features from a reference, guided by geometric priors. Our unified approach operates at 1024px and outperforms specialized methods on tasks ranging from semantic attribute transfer to fine-grained material application. Extensive experiments confirm our state-of-the-art performance in both structural preservation and appearance fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26767
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A training-free framework for high-fidelity appearance transfer via diffusion transformers
Gu, Shengrong
Wang, Ye
Wu, Song
Ma, Rui
Wang, Qian
Wang, Lanjun
Yi, Zili
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first training-free framework specifically designed to tame DiTs for high-fidelity appearance transfer. Our core is a synergistic system that disentangles structure and appearance. We leverage high-fidelity inversion to establish a rich content prior for the source image, capturing its lighting and micro-textures. A novel attention-sharing mechanism then dynamically fuses purified appearance features from a reference, guided by geometric priors. Our unified approach operates at 1024px and outperforms specialized methods on tasks ranging from semantic attribute transfer to fine-grained material application. Extensive experiments confirm our state-of-the-art performance in both structural preservation and appearance fidelity.
title A training-free framework for high-fidelity appearance transfer via diffusion transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.26767