PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shulei, Wei, Longhui, He, Xin, Ouyang, Jianbo, Lu, Hui, Zhao, Zhou, Tian, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024)
by: He, Huiguo, et al.
Published: (2024)
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
by: Cheng, Junhao, et al.
Published: (2024)
by: Cheng, Junhao, et al.
Published: (2024)
Multi-Subject Personalization
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas
by: Yuan, Longhui
Published: (2026)
by: Yuan, Longhui
Published: (2026)
PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
by: Ping, Bowen, et al.
Published: (2025)
by: Ping, Bowen, et al.
Published: (2025)
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
by: Zhou, Yufan, et al.
Published: (2024)
by: Zhou, Yufan, et al.
Published: (2024)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025)
by: Gao, Zhanxin, et al.
Published: (2025)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation
by: He, Junjie, et al.
Published: (2025)
by: He, Junjie, et al.
Published: (2025)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
Identity Decoupling for Multi-Subject Personalization of Text-to-Image Models
by: Jang, Sangwon, et al.
Published: (2024)
by: Jang, Sangwon, et al.
Published: (2024)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
by: Huang, Binyuan, et al.
Published: (2024)
by: Huang, Binyuan, et al.
Published: (2024)
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
by: Ling, Run, et al.
Published: (2025)
by: Ling, Run, et al.
Published: (2025)
Multi-Subject Image Synthesis as a Generative Prior for Single-Subject PET Image Reconstruction
by: Webber, George, et al.
Published: (2024)
by: Webber, George, et al.
Published: (2024)
Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
by: Ma, Jian, et al.
Published: (2023)
by: Ma, Jian, et al.
Published: (2023)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
by: Wei, Tianyi, et al.
Published: (2024)
by: Wei, Tianyi, et al.
Published: (2024)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024)
by: Chan, Kelvin C. K., et al.
Published: (2024)
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
by: Guo, Zinan, et al.
Published: (2025)
by: Guo, Zinan, et al.
Published: (2025)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
by: She, Dong, et al.
Published: (2025)
by: She, Dong, et al.
Published: (2025)
RetriBooru: Leakage-Free Retrieval of Conditions from Reference Images for Subject-Driven Generation
by: Tang, Haoran, et al.
Published: (2023)
by: Tang, Haoran, et al.
Published: (2023)
OVMR: Open-Vocabulary Recognition with Multi-Modal References
by: Ma, Zehong, et al.
Published: (2024)
by: Ma, Zehong, et al.
Published: (2024)
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
Efficient Multi-modal Long Context Learning for Training-free Adaptation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
ComFusion: Personalized Subject Generation in Multiple Specific Scenes From Single Image
by: Hong, Yan, et al.
Published: (2024)
by: Hong, Yan, et al.
Published: (2024)
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
by: Cheng, Yufeng, et al.
Published: (2025)
by: Cheng, Yufeng, et al.
Published: (2025)
MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation
by: Tian, Wenqing, et al.
Published: (2026)
by: Tian, Wenqing, et al.
Published: (2026)
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
by: Chen, Bowen, et al.
Published: (2025)
by: Chen, Bowen, et al.
Published: (2025)
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
by: Jin, Qiaoqiao, et al.
Published: (2025)
by: Jin, Qiaoqiao, et al.
Published: (2025)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
by: Xu, Yijia, et al.
Published: (2026)
by: Xu, Yijia, et al.
Published: (2026)
Geometric Disentanglement of Text Embeddings for Subject-Consistent Text-to-Image Generation using A Single Prompt
by: Li, Shangxun, et al.
Published: (2025)
by: Li, Shangxun, et al.
Published: (2025)
When Identities Collapse: A Stress-Test Benchmark for Multi-Subject Personalization
by: Chen, Zhihan, et al.
Published: (2026)
by: Chen, Zhihan, et al.
Published: (2026)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
FlowFixer: Towards Detail-Preserving Subject-Driven Generation
by: Jun, Jinyoung, et al.
Published: (2026)
by: Jun, Jinyoung, et al.
Published: (2026)
Similar Items
-
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
by: He, Xin, et al.
Published: (2025) -
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024) -
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
by: Cheng, Junhao, et al.
Published: (2024) -
Multi-Subject Personalization
by: Jain, Arushi, et al.
Published: (2024) -
AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas
by: Yuan, Longhui
Published: (2026)