Flux Already Knows -- Activating Subject-Driven Image Generation without Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kang, Hao, Fotiadis, Stathi, Jiang, Liming, Yan, Qing, Jia, Yumin, Liu, Zichuan, Chong, Min Jin, Lu, Xin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915251301122048
author Kang, Hao
Fotiadis, Stathi
Jiang, Liming
Yan, Qing
Jia, Yumin
Liu, Zichuan
Chong, Min Jin
Lu, Xin
author_facet Kang, Hao
Fotiadis, Stathi
Jiang, Liming
Yan, Qing
Jia, Yumin
Liu, Zichuan
Chong, Min Jin
Lu, Xin
contents We propose a simple yet effective zero-shot framework for subject-driven image generation using a vanilla Flux model. By framing the task as grid-based image completion and simply replicating the subject image(s) in a mosaic layout, we activate strong identity-preserving capabilities without any additional data, training, or inference-time fine-tuning. This "free lunch" approach is further strengthened by a novel cascade attention design and meta prompting technique, boosting fidelity and versatility. Experimental results show that our method outperforms baselines across multiple key metrics in benchmarks and human preference studies, with trade-offs in certain aspects. Additionally, it supports diverse edits, including logo insertion, virtual try-on, and subject replacement or insertion. These results demonstrate that a pre-trained foundational text-to-image model can enable high-quality, resource-efficient subject-driven generation, opening new possibilities for lightweight customization in downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flux Already Knows -- Activating Subject-Driven Image Generation without Training
Kang, Hao
Fotiadis, Stathi
Jiang, Liming
Yan, Qing
Jia, Yumin
Liu, Zichuan
Chong, Min Jin
Lu, Xin
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose a simple yet effective zero-shot framework for subject-driven image generation using a vanilla Flux model. By framing the task as grid-based image completion and simply replicating the subject image(s) in a mosaic layout, we activate strong identity-preserving capabilities without any additional data, training, or inference-time fine-tuning. This "free lunch" approach is further strengthened by a novel cascade attention design and meta prompting technique, boosting fidelity and versatility. Experimental results show that our method outperforms baselines across multiple key metrics in benchmarks and human preference studies, with trade-offs in certain aspects. Additionally, it supports diverse edits, including logo insertion, virtual try-on, and subject replacement or insertion. These results demonstrate that a pre-trained foundational text-to-image model can enable high-quality, resource-efficient subject-driven generation, opening new possibilities for lightweight customization in downstream applications.
title Flux Already Knows -- Activating Subject-Driven Image Generation without Training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.11478