FlowFixer: Towards Detail-Preserving Subject-Driven Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jun, Jinyoung, Jang, Won-Dong, Ouyang, Wenbin, Gadde, Raghudeep, Lee, Jungbeom
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911472774283264
author Jun, Jinyoung
Jang, Won-Dong
Ouyang, Wenbin
Gadde, Raghudeep
Lee, Jungbeom
author_facet Jun, Jinyoung
Jang, Won-Dong
Ouyang, Wenbin
Gadde, Raghudeep
Lee, Jungbeom
contents We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21402
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlowFixer: Towards Detail-Preserving Subject-Driven Generation
Jun, Jinyoung
Jang, Won-Dong
Ouyang, Wenbin
Gadde, Raghudeep
Lee, Jungbeom
Computer Vision and Pattern Recognition
We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation.
title FlowFixer: Towards Detail-Preserving Subject-Driven Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.21402