Imagine yourself: Tuning-Free Personalized Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Zecheng, Sun, Bo, Juefei-Xu, Felix, Ma, Haoyu, Ramchandani, Ankit, Cheung, Vincent, Shah, Siddharth, Kalia, Anmol, Subramanyam, Harihar, Zareian, Alireza, Chen, Li, Jain, Ankit, Zhang, Ning, Zhang, Peizhao, Sumbaly, Roshan, Vajda, Peter, Sinha, Animesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916402869305344
author He, Zecheng
Sun, Bo
Juefei-Xu, Felix
Ma, Haoyu
Ramchandani, Ankit
Cheung, Vincent
Shah, Siddharth
Kalia, Anmol
Subramanyam, Harihar
Zareian, Alireza
Chen, Li
Jain, Ankit
Zhang, Ning
Zhang, Peizhao
Sumbaly, Roshan
Vajda, Peter
Sinha, Animesh
author_facet He, Zecheng
Sun, Bo
Juefei-Xu, Felix
Ma, Haoyu
Ramchandani, Ankit
Cheung, Vincent
Shah, Siddharth
Kalia, Anmol
Subramanyam, Harihar
Zareian, Alireza
Chen, Li
Jain, Ankit
Zhang, Ning
Zhang, Peizhao
Sumbaly, Roshan
Vajda, Peter
Sinha, Animesh
contents Diffusion models have demonstrated remarkable efficacy across various image-to-image tasks. In this research, we introduce Imagine yourself, a state-of-the-art model designed for personalized image generation. Unlike conventional tuning-based personalization techniques, Imagine yourself operates as a tuning-free model, enabling all users to leverage a shared framework without individualized adjustments. Moreover, previous work met challenges balancing identity preservation, following complex prompts and preserving good visual quality, resulting in models having strong copy-paste effect of the reference images. Thus, they can hardly generate images following prompts that require significant changes to the reference image, \eg, changing facial expression, head and body poses, and the diversity of the generated images is low. To address these limitations, our proposed method introduces 1) a new synthetic paired data generation mechanism to encourage image diversity, 2) a fully parallel attention architecture with three text encoders and a fully trainable vision encoder to improve the text faithfulness, and 3) a novel coarse-to-fine multi-stage finetuning methodology that gradually pushes the boundary of visual quality. Our study demonstrates that Imagine yourself surpasses the state-of-the-art personalization model, exhibiting superior capabilities in identity preservation, visual quality, and text alignment. This model establishes a robust foundation for various personalization applications. Human evaluation results validate the model's SOTA superiority across all aspects (identity preservation, text faithfulness, and visual appeal) compared to the previous personalization models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13346
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Imagine yourself: Tuning-Free Personalized Image Generation
He, Zecheng
Sun, Bo
Juefei-Xu, Felix
Ma, Haoyu
Ramchandani, Ankit
Cheung, Vincent
Shah, Siddharth
Kalia, Anmol
Subramanyam, Harihar
Zareian, Alireza
Chen, Li
Jain, Ankit
Zhang, Ning
Zhang, Peizhao
Sumbaly, Roshan
Vajda, Peter
Sinha, Animesh
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion models have demonstrated remarkable efficacy across various image-to-image tasks. In this research, we introduce Imagine yourself, a state-of-the-art model designed for personalized image generation. Unlike conventional tuning-based personalization techniques, Imagine yourself operates as a tuning-free model, enabling all users to leverage a shared framework without individualized adjustments. Moreover, previous work met challenges balancing identity preservation, following complex prompts and preserving good visual quality, resulting in models having strong copy-paste effect of the reference images. Thus, they can hardly generate images following prompts that require significant changes to the reference image, \eg, changing facial expression, head and body poses, and the diversity of the generated images is low. To address these limitations, our proposed method introduces 1) a new synthetic paired data generation mechanism to encourage image diversity, 2) a fully parallel attention architecture with three text encoders and a fully trainable vision encoder to improve the text faithfulness, and 3) a novel coarse-to-fine multi-stage finetuning methodology that gradually pushes the boundary of visual quality. Our study demonstrates that Imagine yourself surpasses the state-of-the-art personalization model, exhibiting superior capabilities in identity preservation, visual quality, and text alignment. This model establishes a robust foundation for various personalization applications. Human evaluation results validate the model's SOTA superiority across all aspects (identity preservation, text faithfulness, and visual appeal) compared to the previous personalization models.
title Imagine yourself: Tuning-Free Personalized Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.13346