UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Yufeng, Wu, Wenxu, Wu, Shaojin, Huang, Mengqi, Ding, Fei, He, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
by: Wu, Shaojin, et al.
Published: (2024)
by: Wu, Shaojin, et al.
Published: (2024)
DreamO: A Unified Framework for Image Customization
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)
by: Mao, Zhendong, et al.
Published: (2024)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning
by: Li, Hengjia, et al.
Published: (2025)
by: Li, Hengjia, et al.
Published: (2025)
Stream-T1: Test-Time Scaling for Streaming Video Generation
by: Tu, Yijing, et al.
Published: (2026)
by: Tu, Yijing, et al.
Published: (2026)
PositionIC: Unified Position and Identity Consistency for Image Customization
by: Hu, Junjie, et al.
Published: (2025)
by: Hu, Junjie, et al.
Published: (2025)
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
by: Cong, Xiaoyan, et al.
Published: (2026)
by: Cong, Xiaoyan, et al.
Published: (2026)
Lance: Unified Multimodal Modeling by Multi-Task Synergy
by: Fu, Fengyi, et al.
Published: (2026)
by: Fu, Fengyi, et al.
Published: (2026)
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
by: Wang, Shulei, et al.
Published: (2025)
by: Wang, Shulei, et al.
Published: (2025)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
SCoT: Unifying Consistency Models and Rectified Flows via Straight-Consistent Trajectories
by: Wu, Zhangkai, et al.
Published: (2025)
by: Wu, Zhangkai, et al.
Published: (2025)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
by: Wang, Wenchuan, et al.
Published: (2025)
by: Wang, Wenchuan, et al.
Published: (2025)
RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global Complementation
by: Sun, Zhaoyang, et al.
Published: (2024)
by: Sun, Zhaoyang, et al.
Published: (2024)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
by: Bahng, Hyojin, et al.
Published: (2025)
by: Bahng, Hyojin, et al.
Published: (2025)
Remote Sensing Image Segmentation Using Vision Mamba and Multi-Scale Multi-Frequency Feature Fusion
by: Cao, Yice, et al.
Published: (2024)
by: Cao, Yice, et al.
Published: (2024)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
by: Guo, Zinan, et al.
Published: (2025)
by: Guo, Zinan, et al.
Published: (2025)
Attention Deep Model with Multi-Scale Deep Supervision for Person Re-Identification
by: Wu, Di, et al.
Published: (2019)
by: Wu, Di, et al.
Published: (2019)
Generating Multi-Image Synthetic Data for Text-to-Image Customization
by: Kumari, Nupur, et al.
Published: (2025)
by: Kumari, Nupur, et al.
Published: (2025)
FLUX-Makeup: High-Fidelity, Identity-Consistent, and Robust Makeup Transfer via Diffusion Transformer
by: Zhu, Jian, et al.
Published: (2025)
by: Zhu, Jian, et al.
Published: (2025)
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
by: Chen, Bowen, et al.
Published: (2025)
by: Chen, Bowen, et al.
Published: (2025)
ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
by: Wu, Mingyang, et al.
Published: (2026)
by: Wu, Mingyang, et al.
Published: (2026)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024)
by: Oertell, Owen, et al.
Published: (2024)
GC-MVSNet: Multi-View, Multi-Scale, Geometrically-Consistent Multi-View Stereo
by: Vats, Vibhas K., et al.
Published: (2023)
by: Vats, Vibhas K., et al.
Published: (2023)
CustomText: Customized Textual Image Generation using Diffusion Models
by: Paliwal, Shubham, et al.
Published: (2024)
by: Paliwal, Shubham, et al.
Published: (2024)
OwMatch: Conditional Self-Labeling with Consistency for Open-World Semi-Supervised Learning
by: Niu, Shengjie, et al.
Published: (2024)
by: Niu, Shengjie, et al.
Published: (2024)
Temporally Consistent Stereo Matching
by: Zeng, Jiaxi, et al.
Published: (2024)
by: Zeng, Jiaxi, et al.
Published: (2024)
A Robust Multisource Remote Sensing Image Matching Method Utilizing Attention and Feature Enhancement Against Noise Interference
by: Li, Yuan, et al.
Published: (2024)
by: Li, Yuan, et al.
Published: (2024)
Enhancing Diffusion-Based Quantitatively Controllable Image Generation via Matrix-Form EDM and Adaptive Vicinal Training
by: Ding, Xin, et al.
Published: (2026)
by: Ding, Xin, et al.
Published: (2026)
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
by: Shen, Liao, et al.
Published: (2025)
by: Shen, Liao, et al.
Published: (2025)
PuLID: Pure and Lightning ID Customization via Contrastive Alignment
by: Guo, Zinan, et al.
Published: (2024)
by: Guo, Zinan, et al.
Published: (2024)
PIDiff: Image Customization for Personalized Identities with Diffusion Models
by: Gu, Jinyu, et al.
Published: (2025)
by: Gu, Jinyu, et al.
Published: (2025)
Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation
by: Ye, Jinyan, et al.
Published: (2026)
by: Ye, Jinyan, et al.
Published: (2026)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023)
by: Huang, Tao, et al.
Published: (2023)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
by: Bai, Shaojin, et al.
Published: (2026)
by: Bai, Shaojin, et al.
Published: (2026)
Similar Items
-
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
by: Wu, Shaojin, et al.
Published: (2025) -
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025) -
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
by: Wu, Shaojin, et al.
Published: (2024) -
DreamO: A Unified Framework for Image Customization
by: Mou, Chong, et al.
Published: (2025) -
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)