Unified Personalized Understanding, Generating and Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Yu, Lin, Tianwei, Zhu, Ruike, Yuan, Yuqian, Zheng, Haoyu, Liang, Liang, Zhang, Wenqiao, Shao, Feifei, Li, Haoyuan, He, Wanggui, Jiang, Hao, Zhuang, Yueting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911368220770304
author Zhong, Yu
Lin, Tianwei
Zhu, Ruike
Yuan, Yuqian
Zheng, Haoyu
Liang, Liang
Zhang, Wenqiao
Shao, Feifei
Li, Haoyuan
He, Wanggui
Jiang, Hao
Zhuang, Yueting
author_facet Zhong, Yu
Lin, Tianwei
Zhu, Ruike
Yuan, Yuqian
Zheng, Haoyu
Liang, Liang
Zhang, Wenqiao
Shao, Feifei
Li, Haoyuan
He, Wanggui
Jiang, Hao
Zhuang, Yueting
contents Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific concepts (e.g., generate a photo of \texttt{<maeve>}) in a consistent and controllable manner. Existing personalization methods typically rely on external retrieval, which is inefficient and poorly integrated into unified multimodal pipelines. Recent personalized unified models introduce learnable soft prompts to encode concept information, yet they either couple understanding and generation or depend on complex multi-stage training, leading to cross-task interference and ultimately to fuzzy or misaligned personalized knowledge. We present \textbf{OmniPersona}, an end-to-end personalization framework for unified LMMs that, for the first time, integrates personalized understanding, generation, and image editing within a single architecture. OmniPersona introduces structurally decoupled concept tokens, allocating dedicated subspaces for different tasks to minimize interference, and incorporates an explicit knowledge replay mechanism that propagates personalized attribute knowledge across tasks, enabling consistent personalized behavior. To systematically evaluate unified personalization, we propose \textbf{\texttt{OmniPBench}}, extending the public UnifyBench concept set with personalized editing tasks and cross-task evaluation protocols integrating understanding, generation, and editing. Experimental results demonstrate that OmniPersona delivers competitive and robust performance across diverse personalization tasks. We hope OmniPersona will serve as a strong baseline and spur further research on controllable, unified personalization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06965
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unified Personalized Understanding, Generating and Editing
Zhong, Yu
Lin, Tianwei
Zhu, Ruike
Yuan, Yuqian
Zheng, Haoyu
Liang, Liang
Zhang, Wenqiao
Shao, Feifei
Li, Haoyuan
He, Wanggui
Jiang, Hao
Zhuang, Yueting
Computer Vision and Pattern Recognition
Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific concepts (e.g., generate a photo of \texttt{<maeve>}) in a consistent and controllable manner. Existing personalization methods typically rely on external retrieval, which is inefficient and poorly integrated into unified multimodal pipelines. Recent personalized unified models introduce learnable soft prompts to encode concept information, yet they either couple understanding and generation or depend on complex multi-stage training, leading to cross-task interference and ultimately to fuzzy or misaligned personalized knowledge. We present \textbf{OmniPersona}, an end-to-end personalization framework for unified LMMs that, for the first time, integrates personalized understanding, generation, and image editing within a single architecture. OmniPersona introduces structurally decoupled concept tokens, allocating dedicated subspaces for different tasks to minimize interference, and incorporates an explicit knowledge replay mechanism that propagates personalized attribute knowledge across tasks, enabling consistent personalized behavior. To systematically evaluate unified personalization, we propose \textbf{\texttt{OmniPBench}}, extending the public UnifyBench concept set with personalized editing tasks and cross-task evaluation protocols integrating understanding, generation, and editing. Experimental results demonstrate that OmniPersona delivers competitive and robust performance across diverse personalization tasks. We hope OmniPersona will serve as a strong baseline and spur further research on controllable, unified personalization.
title Unified Personalized Understanding, Generating and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.06965