A Survey on Personalized Content Synthesis with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xulu, Wei, Xiaoyong, Hu, Wentao, Wu, Jinlin, Wu, Jiaxin, Zhang, Wengyu, Zhang, Zhaoxiang, Lei, Zhen, Li, Qing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914172012331008
author Zhang, Xulu
Wei, Xiaoyong
Hu, Wentao
Wu, Jinlin
Wu, Jiaxin
Zhang, Wengyu
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
author_facet Zhang, Xulu
Wei, Xiaoyong
Hu, Wentao
Wu, Jinlin
Wu, Jiaxin
Zhang, Wengyu
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
contents Recent advancements in diffusion models have significantly impacted content creation, leading to the emergence of Personalized Content Synthesis (PCS). By utilizing a small set of user-provided examples featuring the same subject, PCS aims to tailor this subject to specific user-defined prompts. Over the past two years, more than 150 methods have been introduced in this area. However, existing surveys primarily focus on text-to-image generation, with few providing up-to-date summaries on PCS. This paper provides a comprehensive survey of PCS, introducing the general frameworks of PCS research, which can be categorized into test-time fine-tuning (TTF) and pre-trained adaptation (PTA) approaches. We analyze the strengths, limitations, and key techniques of these methodologies. Additionally, we explore specialized tasks within the field, such as object, face, and style personalization, while highlighting their unique challenges and innovations. Despite the promising progress, we also discuss ongoing challenges, including overfitting and the trade-off between subject fidelity and text alignment. Through this detailed overview and analysis, we propose future directions to further the development of PCS.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey on Personalized Content Synthesis with Diffusion Models
Zhang, Xulu
Wei, Xiaoyong
Hu, Wentao
Wu, Jinlin
Wu, Jiaxin
Zhang, Wengyu
Zhang, Zhaoxiang
Lei, Zhen
Li, Qing
Computer Vision and Pattern Recognition
Recent advancements in diffusion models have significantly impacted content creation, leading to the emergence of Personalized Content Synthesis (PCS). By utilizing a small set of user-provided examples featuring the same subject, PCS aims to tailor this subject to specific user-defined prompts. Over the past two years, more than 150 methods have been introduced in this area. However, existing surveys primarily focus on text-to-image generation, with few providing up-to-date summaries on PCS. This paper provides a comprehensive survey of PCS, introducing the general frameworks of PCS research, which can be categorized into test-time fine-tuning (TTF) and pre-trained adaptation (PTA) approaches. We analyze the strengths, limitations, and key techniques of these methodologies. Additionally, we explore specialized tasks within the field, such as object, face, and style personalization, while highlighting their unique challenges and innovations. Despite the promising progress, we also discuss ongoing challenges, including overfitting and the trade-off between subject fidelity and text alignment. Through this detailed overview and analysis, we propose future directions to further the development of PCS.
title A Survey on Personalized Content Synthesis with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.05538