UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yiyan, Wang, Qiulin, Wang, Wenjie, Mao, Yunyao, Wang, Xintao, Wan, Pengfei, Gai, Kun, Feng, Fuli |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniVideo: Unified Understanding, Generation, and Editing for Videos
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025)
by: Huang, Yuzhou, et al.
Published: (2025)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
by: Cai, Minghong, et al.
Published: (2025)
by: Cai, Minghong, et al.
Published: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
UNIC: Unified In-Context Video Editing
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
by: Luo, Yawen, et al.
Published: (2025)
by: Luo, Yawen, et al.
Published: (2025)
RelightMaster: Precise Video Relighting with Multi-plane Light Images
by: Bian, Weikang, et al.
Published: (2025)
by: Bian, Weikang, et al.
Published: (2025)
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
by: Wang, Qinghe, et al.
Published: (2025)
by: Wang, Qinghe, et al.
Published: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual Correspondence
by: Wu, Ruihai, et al.
Published: (2024)
by: Wu, Ruihai, et al.
Published: (2024)
DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
by: Xu, Yiyan, et al.
Published: (2025)
by: Xu, Yiyan, et al.
Published: (2025)
NextAds: Towards Next-generation Personalized Video Advertising
by: Xu, Yiyan, et al.
Published: (2026)
by: Xu, Yiyan, et al.
Published: (2026)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
by: Wang, Qinghe, et al.
Published: (2025)
by: Wang, Qinghe, et al.
Published: (2025)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
by: Zhou, Zhaokun, et al.
Published: (2024)
by: Zhou, Zhaokun, et al.
Published: (2024)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
Uni-Animator: Towards Unified Visual Colorization
by: Chen, Xinyuan, et al.
Published: (2026)
by: Chen, Xinyuan, et al.
Published: (2026)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
A Survey of Interactive Generative Video
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
by: Wang, Haoxuan, et al.
Published: (2025)
by: Wang, Haoxuan, et al.
Published: (2025)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
MultiRef: Controllable Image Generation with Multiple Visual References
by: Chen, Ruoxi, et al.
Published: (2025)
by: Chen, Ruoxi, et al.
Published: (2025)
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
by: Wang, Yuanrui, et al.
Published: (2025)
by: Wang, Yuanrui, et al.
Published: (2025)
Conditional Text-to-Image Generation with Reference Guidance
by: Kim, Taewook, et al.
Published: (2024)
by: Kim, Taewook, et al.
Published: (2024)
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
by: Liao, Zhichao, et al.
Published: (2025)
by: Liao, Zhichao, et al.
Published: (2025)
StyleAdapter: A Unified Stylized Image Generation Model
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
SemanticGen: Video Generation in Semantic Space
by: Bai, Jianhong, et al.
Published: (2025)
by: Bai, Jianhong, et al.
Published: (2025)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Personalized Image Generation with Large Multimodal Models
by: Xu, Yiyan, et al.
Published: (2024)
by: Xu, Yiyan, et al.
Published: (2024)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
by: Wang, Jiayun, et al.
Published: (2026)
by: Wang, Jiayun, et al.
Published: (2026)
UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
by: Cui, Haowang, et al.
Published: (2025)
by: Cui, Haowang, et al.
Published: (2025)
Geometry-Aware Implicit Memory for Video World Models
by: Wei, Zhengxuan, et al.
Published: (2026)
by: Wei, Zhengxuan, et al.
Published: (2026)
Similar Items
-
UniVideo: Unified Understanding, Generation, and Editing for Videos
by: Wei, Cong, et al.
Published: (2025) -
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025) -
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
by: Cai, Minghong, et al.
Published: (2025) -
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025) -
UNIC: Unified In-Context Video Editing
by: Ye, Zixuan, et al.
Published: (2025)