USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shaojin, Huang, Mengqi, Cheng, Yufeng, Wu, Wenxu, Tian, Jiahe, Luo, Yiming, Ding, Fei, He, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
by: Cheng, Yufeng, et al.
Published: (2025)
by: Cheng, Yufeng, et al.
Published: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
by: Wu, Shaojin, et al.
Published: (2024)
by: Wu, Shaojin, et al.
Published: (2024)
DreamO: A Unified Framework for Image Customization
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
Lance: Unified Multimodal Modeling by Multi-Task Synergy
by: Fu, Fengyi, et al.
Published: (2026)
by: Fu, Fengyi, et al.
Published: (2026)
UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
by: Yang, Jingwei, et al.
Published: (2026)
by: Yang, Jingwei, et al.
Published: (2026)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
by: Pan, Jiadong, et al.
Published: (2026)
by: Pan, Jiadong, et al.
Published: (2026)
Stream-T1: Test-Time Scaling for Streaming Video Generation
by: Tu, Yijing, et al.
Published: (2026)
by: Tu, Yijing, et al.
Published: (2026)
Disentangled Latent Energy-Based Style Translation: An Image-Level Structural MRI Harmonization Framework
by: Wu, Mengqi, et al.
Published: (2024)
by: Wu, Mengqi, et al.
Published: (2024)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints
by: Wu, Yuxuan, et al.
Published: (2024)
by: Wu, Yuxuan, et al.
Published: (2024)
Unified Multi-Site Multi-Sequence Brain MRI Harmonization Enriched by Biomedical Semantic Style
by: Wu, Mengqi, et al.
Published: (2026)
by: Wu, Mengqi, et al.
Published: (2026)
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
by: Wang, Shulei, et al.
Published: (2025)
by: Wang, Shulei, et al.
Published: (2025)
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
by: Gong, Yuan, et al.
Published: (2025)
by: Gong, Yuan, et al.
Published: (2025)
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
by: Fu, Fengyi, et al.
Published: (2025)
by: Fu, Fengyi, et al.
Published: (2025)
CDST: Color Disentangled Style Transfer for Universal Style Reference Customization
by: Zhang, Shiwen, et al.
Published: (2025)
by: Zhang, Shiwen, et al.
Published: (2025)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization
by: Zhao, Boyang, et al.
Published: (2026)
by: Zhao, Boyang, et al.
Published: (2026)
ConDSeg: A General Medical Image Segmentation Framework via Contrast-Driven Feature Enhancement
by: Lei, Mengqi, et al.
Published: (2024)
by: Lei, Mengqi, et al.
Published: (2024)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion
by: Peng, Fei, et al.
Published: (2025)
by: Peng, Fei, et al.
Published: (2025)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
by: Chen, Qiyuan, et al.
Published: (2026)
by: Chen, Qiyuan, et al.
Published: (2026)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)
by: Mao, Zhendong, et al.
Published: (2024)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
by: Tian, Juanxi, et al.
Published: (2025)
by: Tian, Juanxi, et al.
Published: (2025)
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
by: Guo, Zinan, et al.
Published: (2025)
by: Guo, Zinan, et al.
Published: (2025)
PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
StyleDecoupler: Generalizable Artistic Style Disentanglement
by: Jia, Zexi, et al.
Published: (2026)
by: Jia, Zexi, et al.
Published: (2026)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
by: Bai, Shaojin, et al.
Published: (2026)
by: Bai, Shaojin, et al.
Published: (2026)
I2VControl: Disentangled and Unified Video Motion Synthesis Control
by: Feng, Wanquan, et al.
Published: (2024)
by: Feng, Wanquan, et al.
Published: (2024)
Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
by: Li, Shuang, et al.
Published: (2026)
by: Li, Shuang, et al.
Published: (2026)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
WikiStyle+: A Multimodal Approach to Content-Style Representation Disentanglement for Artistic Image Stylization
by: Zhuoqi, Ma, et al.
Published: (2024)
by: Zhuoqi, Ma, et al.
Published: (2024)
DreamStyle: A Unified Framework for Video Stylization
by: Li, Mengtian, et al.
Published: (2026)
by: Li, Mengtian, et al.
Published: (2026)
Similar Items
-
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
by: Cheng, Yufeng, et al.
Published: (2025) -
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025) -
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
by: Wu, Shaojin, et al.
Published: (2024) -
DreamO: A Unified Framework for Image Customization
by: Mou, Chong, et al.
Published: (2025) -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)