MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xujie, Lin, Ente, Li, Xiu, Luo, Yuxuan, Kampffmeyer, Michael, Dong, Xin, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
von: Lin, Ente, et al.
Veröffentlicht: (2024)
von: Lin, Ente, et al.
Veröffentlicht: (2024)
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
von: Xie, Zhenyu, et al.
Veröffentlicht: (2026)
von: Xie, Zhenyu, et al.
Veröffentlicht: (2026)
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
von: Chen, Jianqi, et al.
Veröffentlicht: (2024)
von: Chen, Jianqi, et al.
Veröffentlicht: (2024)
ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions
von: Zhang, Shiyue, et al.
Veröffentlicht: (2025)
von: Zhang, Shiyue, et al.
Veröffentlicht: (2025)
ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting
von: Ma, Zehua, et al.
Veröffentlicht: (2026)
von: Ma, Zehua, et al.
Veröffentlicht: (2026)
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
von: Kong, Xianghao, et al.
Veröffentlicht: (2025)
von: Kong, Xianghao, et al.
Veröffentlicht: (2025)
FashionComposer: Compositional Fashion Image Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
von: Ding, Xinpeng, et al.
Veröffentlicht: (2024)
von: Ding, Xinpeng, et al.
Veröffentlicht: (2024)
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
OVMR: Open-Vocabulary Recognition with Multi-Modal References
von: Ma, Zehong, et al.
Veröffentlicht: (2024)
von: Ma, Zehong, et al.
Veröffentlicht: (2024)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
von: Zuo, Tongchun, et al.
Veröffentlicht: (2025)
von: Zuo, Tongchun, et al.
Veröffentlicht: (2025)
Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models
von: Lv, Weiyi, et al.
Veröffentlicht: (2025)
von: Lv, Weiyi, et al.
Veröffentlicht: (2025)
IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion Design
von: Shen, Fei, et al.
Veröffentlicht: (2025)
von: Shen, Fei, et al.
Veröffentlicht: (2025)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
von: Liao, Chao, et al.
Veröffentlicht: (2025)
von: Liao, Chao, et al.
Veröffentlicht: (2025)
Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
von: Huang, Jiancheng, et al.
Veröffentlicht: (2024)
von: Huang, Jiancheng, et al.
Veröffentlicht: (2024)
Quality and Quantity: Unveiling a Million High-Quality Images for Text-to-Image Synthesis in Fashion Design
von: Yu, Jia, et al.
Veröffentlicht: (2023)
von: Yu, Jia, et al.
Veröffentlicht: (2023)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
von: Pan, Wensheng, et al.
Veröffentlicht: (2024)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
von: Peng, Xinge, et al.
Veröffentlicht: (2026)
von: Peng, Xinge, et al.
Veröffentlicht: (2026)
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
von: Liu, Keli, et al.
Veröffentlicht: (2025)
von: Liu, Keli, et al.
Veröffentlicht: (2025)
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
von: Zhang, Rong, et al.
Veröffentlicht: (2025)
von: Zhang, Rong, et al.
Veröffentlicht: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
von: Huang, Binyuan, et al.
Veröffentlicht: (2026)
von: Huang, Binyuan, et al.
Veröffentlicht: (2026)
M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension
von: Liu, Xuyang, et al.
Veröffentlicht: (2024)
von: Liu, Xuyang, et al.
Veröffentlicht: (2024)
Multi-Modal Generative Embedding Model
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval
von: Pan, Yi, et al.
Veröffentlicht: (2025)
von: Pan, Yi, et al.
Veröffentlicht: (2025)
ControlEdit: A MultiModal Local Clothing Image Editing Method
von: Cheng, Di, et al.
Veröffentlicht: (2024)
von: Cheng, Di, et al.
Veröffentlicht: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
von: Li, Liyang, et al.
Veröffentlicht: (2026)
von: Li, Liyang, et al.
Veröffentlicht: (2026)
User-Friendly Customized Generation with Multi-Modal Prompts
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
von: Lin, Ente, et al.
Veröffentlicht: (2024) -
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
von: Xie, Zhenyu, et al.
Veröffentlicht: (2026) -
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024) -
Collaborative Multi-Modal Coding for High-Quality 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025) -
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
von: Chen, Jianqi, et al.
Veröffentlicht: (2024)