Preventing Shortcuts in Adapter Training via Providing the Shortcuts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goyal, Anujraaj Argo, Qian, Guocheng Gordon, Coskun, Huseyin, Gupta, Aarush, Tam, Himmy, Ostashev, Daniil, Hu, Ju, Sagar, Dhritiman, Tulyakov, Sergey, Aberman, Kfir, Wang, Kuan-Chieh Jackson
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917038717403136
author Goyal, Anujraaj Argo
Qian, Guocheng Gordon
Coskun, Huseyin
Gupta, Aarush
Tam, Himmy
Ostashev, Daniil
Hu, Ju
Sagar, Dhritiman
Tulyakov, Sergey
Aberman, Kfir
Wang, Kuan-Chieh Jackson
author_facet Goyal, Anujraaj Argo
Qian, Guocheng Gordon
Coskun, Huseyin
Gupta, Aarush
Tam, Himmy
Ostashev, Daniil
Hu, Ju
Sagar, Dhritiman
Tulyakov, Sergey
Aberman, Kfir
Wang, Kuan-Chieh Jackson
contents Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synthesis. These adapters are typically trained to capture a specific target attribute, such as subject identity, using single-image reconstruction objectives. However, because the input image inevitably contains a mixture of visual factors, adapters are prone to entangle the target attribute with incidental ones, such as pose, expression, and lighting. This spurious correlation problem limits generalization and obstructs the model's ability to adhere to the input text prompt. In this work, we uncover a simple yet effective solution: provide the very shortcuts we wish to eliminate during adapter training. In Shortcut-Rerouted Adapter Training, confounding factors are routed through auxiliary modules, such as ControlNet or LoRA, eliminating the incentive for the adapter to internalize them. The auxiliary modules are then removed during inference. When applied to tasks like facial and full-body identity injection, our approach improves generation quality, diversity, and prompt adherence. These results point to a general design principle in the era of large models: when seeking disentangled representations, the most effective path may be to establish shortcuts for what should NOT be learned.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preventing Shortcuts in Adapter Training via Providing the Shortcuts
Goyal, Anujraaj Argo
Qian, Guocheng Gordon
Coskun, Huseyin
Gupta, Aarush
Tam, Himmy
Ostashev, Daniil
Hu, Ju
Sagar, Dhritiman
Tulyakov, Sergey
Aberman, Kfir
Wang, Kuan-Chieh Jackson
Computer Vision and Pattern Recognition
Artificial Intelligence
Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synthesis. These adapters are typically trained to capture a specific target attribute, such as subject identity, using single-image reconstruction objectives. However, because the input image inevitably contains a mixture of visual factors, adapters are prone to entangle the target attribute with incidental ones, such as pose, expression, and lighting. This spurious correlation problem limits generalization and obstructs the model's ability to adhere to the input text prompt. In this work, we uncover a simple yet effective solution: provide the very shortcuts we wish to eliminate during adapter training. In Shortcut-Rerouted Adapter Training, confounding factors are routed through auxiliary modules, such as ControlNet or LoRA, eliminating the incentive for the adapter to internalize them. The auxiliary modules are then removed during inference. When applied to tasks like facial and full-body identity injection, our approach improves generation quality, diversity, and prompt adherence. These results point to a general design principle in the era of large models: when seeking disentangled representations, the most effective path may be to establish shortcuts for what should NOT be learned.
title Preventing Shortcuts in Adapter Training via Providing the Shortcuts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.20887