FlowFixer: Towards Detail-Preserving Subject-Driven Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911472774283264 |
|---|---|
| author | Jun, Jinyoung Jang, Won-Dong Ouyang, Wenbin Gadde, Raghudeep Lee, Jungbeom |
| author_facet | Jun, Jinyoung Jang, Won-Dong Ouyang, Wenbin Gadde, Raghudeep Lee, Jungbeom |
| contents | We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_21402 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FlowFixer: Towards Detail-Preserving Subject-Driven Generation Jun, Jinyoung Jang, Won-Dong Ouyang, Wenbin Gadde, Raghudeep Lee, Jungbeom Computer Vision and Pattern Recognition We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation. |
| title | FlowFixer: Towards Detail-Preserving Subject-Driven Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.21402 |