Does FLUX Already Know How to Perform Physically Plausible Image Composition?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lu, Shilin, Lian, Zhuming, Zhou, Zihan, Zhang, Shaocong, Zhao, Chen, Kong, Adams Wai-Kin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912931153707008
author Lu, Shilin
Lian, Zhuming
Zhou, Zihan
Zhang, Shaocong
Zhao, Chen
Kong, Adams Wai-Kin
author_facet Lu, Shilin
Lian, Zhuming
Zhou, Zihan
Zhang, Shaocong
Zhao, Chen
Kong, Adams Wai-Kin
contents Image composition aims to seamlessly insert a user-specified object into a new scene, but existing models struggle with complex lighting (e.g., accurate shadows, water reflections) and diverse, high-resolution inputs. Modern text-to-image diffusion models (e.g., SD3.5, FLUX) already encode essential physical and resolution priors, yet lack a framework to unleash them without resorting to latent inversion, which often locks object poses into contextually inappropriate orientations, or brittle attention surgery. We propose SHINE, a training-free framework for Seamless, High-fidelity Insertion with Neutralized Errors. SHINE introduces manifold-steered anchor loss, leveraging pretrained customization adapters (e.g., IP-Adapter) to guide latents for faithful subject representation while preserving background integrity. Degradation-suppression guidance and adaptive background blending are proposed to further eliminate low-quality outputs and visible seams. To address the lack of rigorous benchmarks, we introduce ComplexCompo, featuring diverse resolutions and challenging conditions such as low lighting, strong illumination, intricate shadows, and reflective surfaces. Experiments on ComplexCompo and DreamEditBench show state-of-the-art performance on standard metrics (e.g., DINOv2) and human-aligned scores (e.g., DreamSim, ImageReward, VisionReward). Code is available at https://github.com/ZhumingLian/SHINE.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does FLUX Already Know How to Perform Physically Plausible Image Composition?
Lu, Shilin
Lian, Zhuming
Zhou, Zihan
Zhang, Shaocong
Zhao, Chen
Kong, Adams Wai-Kin
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image composition aims to seamlessly insert a user-specified object into a new scene, but existing models struggle with complex lighting (e.g., accurate shadows, water reflections) and diverse, high-resolution inputs. Modern text-to-image diffusion models (e.g., SD3.5, FLUX) already encode essential physical and resolution priors, yet lack a framework to unleash them without resorting to latent inversion, which often locks object poses into contextually inappropriate orientations, or brittle attention surgery. We propose SHINE, a training-free framework for Seamless, High-fidelity Insertion with Neutralized Errors. SHINE introduces manifold-steered anchor loss, leveraging pretrained customization adapters (e.g., IP-Adapter) to guide latents for faithful subject representation while preserving background integrity. Degradation-suppression guidance and adaptive background blending are proposed to further eliminate low-quality outputs and visible seams. To address the lack of rigorous benchmarks, we introduce ComplexCompo, featuring diverse resolutions and challenging conditions such as low lighting, strong illumination, intricate shadows, and reflective surfaces. Experiments on ComplexCompo and DreamEditBench show state-of-the-art performance on standard metrics (e.g., DINOv2) and human-aligned scores (e.g., DreamSim, ImageReward, VisionReward). Code is available at https://github.com/ZhumingLian/SHINE.
title Does FLUX Already Know How to Perform Physically Plausible Image Composition?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.21278