MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xujie, Lin, Ente, Li, Xiu, Luo, Yuxuan, Kampffmeyer, Michael, Dong, Xin, Liang, Xiaodan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
di: Lin, Ente, et al.
Pubblicazione: (2024)
di: Lin, Ente, et al.
Pubblicazione: (2024)
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
di: Xie, Zhenyu, et al.
Pubblicazione: (2026)
di: Xie, Zhenyu, et al.
Pubblicazione: (2026)
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
di: Zhang, Shiyue, et al.
Pubblicazione: (2024)
di: Zhang, Shiyue, et al.
Pubblicazione: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
di: Cao, Ziang, et al.
Pubblicazione: (2025)
di: Cao, Ziang, et al.
Pubblicazione: (2025)
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
di: Chen, Jianqi, et al.
Pubblicazione: (2024)
di: Chen, Jianqi, et al.
Pubblicazione: (2024)
ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions
di: Zhang, Shiyue, et al.
Pubblicazione: (2025)
di: Zhang, Shiyue, et al.
Pubblicazione: (2025)
ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting
di: Ma, Zehua, et al.
Pubblicazione: (2026)
di: Ma, Zehua, et al.
Pubblicazione: (2026)
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
di: Zheng, Jun, et al.
Pubblicazione: (2024)
di: Zheng, Jun, et al.
Pubblicazione: (2024)
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
di: Kong, Xianghao, et al.
Pubblicazione: (2025)
di: Kong, Xianghao, et al.
Pubblicazione: (2025)
FashionComposer: Compositional Fashion Image Generation
di: Ji, Sihui, et al.
Pubblicazione: (2024)
di: Ji, Sihui, et al.
Pubblicazione: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
di: Ding, Xinpeng, et al.
Pubblicazione: (2024)
di: Ding, Xinpeng, et al.
Pubblicazione: (2024)
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
di: Chong, Zheng, et al.
Pubblicazione: (2025)
di: Chong, Zheng, et al.
Pubblicazione: (2025)
OVMR: Open-Vocabulary Recognition with Multi-Modal References
di: Ma, Zehong, et al.
Pubblicazione: (2024)
di: Ma, Zehong, et al.
Pubblicazione: (2024)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
di: He, Yu, et al.
Pubblicazione: (2026)
di: He, Yu, et al.
Pubblicazione: (2026)
DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
di: Zuo, Tongchun, et al.
Pubblicazione: (2025)
di: Zuo, Tongchun, et al.
Pubblicazione: (2025)
Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models
di: Lv, Weiyi, et al.
Pubblicazione: (2025)
di: Lv, Weiyi, et al.
Pubblicazione: (2025)
IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion Design
di: Shen, Fei, et al.
Pubblicazione: (2025)
di: Shen, Fei, et al.
Pubblicazione: (2025)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
di: Liao, Chao, et al.
Pubblicazione: (2025)
di: Liao, Chao, et al.
Pubblicazione: (2025)
Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition
di: Liang, Siyu, et al.
Pubblicazione: (2025)
di: Liang, Siyu, et al.
Pubblicazione: (2025)
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
di: Huang, Jiancheng, et al.
Pubblicazione: (2024)
di: Huang, Jiancheng, et al.
Pubblicazione: (2024)
Quality and Quantity: Unveiling a Million High-Quality Images for Text-to-Image Synthesis in Fashion Design
di: Yu, Jia, et al.
Pubblicazione: (2023)
di: Yu, Jia, et al.
Pubblicazione: (2023)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
di: Pan, Wensheng, et al.
Pubblicazione: (2024)
di: Pan, Wensheng, et al.
Pubblicazione: (2024)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
di: Peng, Xinge, et al.
Pubblicazione: (2026)
di: Peng, Xinge, et al.
Pubblicazione: (2026)
MultiRef: Controllable Image Generation with Multiple Visual References
di: Chen, Ruoxi, et al.
Pubblicazione: (2025)
di: Chen, Ruoxi, et al.
Pubblicazione: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
di: Yuan, Peng, et al.
Pubblicazione: (2026)
di: Yuan, Peng, et al.
Pubblicazione: (2026)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
di: Yang, Zhengwei, et al.
Pubblicazione: (2026)
di: Yang, Zhengwei, et al.
Pubblicazione: (2026)
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
di: Liu, Ziyue, et al.
Pubblicazione: (2026)
di: Liu, Ziyue, et al.
Pubblicazione: (2026)
ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
di: Liu, Keli, et al.
Pubblicazione: (2025)
di: Liu, Keli, et al.
Pubblicazione: (2025)
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
di: Zhang, Rong, et al.
Pubblicazione: (2025)
di: Zhang, Rong, et al.
Pubblicazione: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
di: Huang, Binyuan, et al.
Pubblicazione: (2026)
di: Huang, Binyuan, et al.
Pubblicazione: (2026)
M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension
di: Liu, Xuyang, et al.
Pubblicazione: (2024)
di: Liu, Xuyang, et al.
Pubblicazione: (2024)
Multi-Modal Generative Embedding Model
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval
di: Pan, Yi, et al.
Pubblicazione: (2025)
di: Pan, Yi, et al.
Pubblicazione: (2025)
ControlEdit: A MultiModal Local Clothing Image Editing Method
di: Cheng, Di, et al.
Pubblicazione: (2024)
di: Cheng, Di, et al.
Pubblicazione: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
User-Friendly Customized Generation with Multi-Modal Prompts
di: Zhong, Linhao, et al.
Pubblicazione: (2024)
di: Zhong, Linhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
di: Lin, Ente, et al.
Pubblicazione: (2024) -
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
di: Xie, Zhenyu, et al.
Pubblicazione: (2026) -
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
di: Zhang, Shiyue, et al.
Pubblicazione: (2024) -
Collaborative Multi-Modal Coding for High-Quality 3D Generation
di: Cao, Ziang, et al.
Pubblicazione: (2025) -
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
di: Chen, Jianqi, et al.
Pubblicazione: (2024)