UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xiangyu, Zhang, Yuehan, Zhang, Wenlong, Wu, Xiao-Ming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
by: Pang, Kaicheng, et al.
Published: (2025)
by: Pang, Kaicheng, et al.
Published: (2025)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)
by: Singh, Abhishek Kumar, et al.
Published: (2024)
FashionFlow: Leveraging Diffusion Models for Dynamic Fashion Video Synthesis from Static Imagery
by: Islam, Tasin, et al.
Published: (2023)
by: Islam, Tasin, et al.
Published: (2023)
FashionFail: Addressing Failure Cases in Fashion Object Detection and Segmentation
by: Velioglu, Riza, et al.
Published: (2024)
by: Velioglu, Riza, et al.
Published: (2024)
DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing
by: Wang, Xiaolong, et al.
Published: (2024)
by: Wang, Xiaolong, et al.
Published: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
by: Gondal, Moazzam Umer, et al.
Published: (2025)
by: Gondal, Moazzam Umer, et al.
Published: (2025)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
by: Choi, Jinyoung, et al.
Published: (2025)
by: Choi, Jinyoung, et al.
Published: (2025)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection
by: Sun, Fuxiang, et al.
Published: (2026)
by: Sun, Fuxiang, et al.
Published: (2026)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
by: Li, Jinke, et al.
Published: (2025)
by: Li, Jinke, et al.
Published: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
FEAT: Fashion Editing and Try-On from Any Design
by: Kwon, Soye, et al.
Published: (2026)
by: Kwon, Soye, et al.
Published: (2026)
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
by: Xiao, Ling, et al.
Published: (2022)
by: Xiao, Ling, et al.
Published: (2022)
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
by: Girella, Federico, et al.
Published: (2025)
by: Girella, Federico, et al.
Published: (2025)
FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
by: Wen, Haokun, et al.
Published: (2026)
by: Wen, Haokun, et al.
Published: (2026)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion Models
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
by: Wang, Zhen, et al.
Published: (2023)
by: Wang, Zhen, et al.
Published: (2023)
UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization
by: Huang, Qing, et al.
Published: (2025)
by: Huang, Qing, et al.
Published: (2025)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
by: Han, Ruiyan, et al.
Published: (2026)
by: Han, Ruiyan, et al.
Published: (2026)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
by: Huang, Jiale, et al.
Published: (2024)
by: Huang, Jiale, et al.
Published: (2024)
UniBEVFusion: Unified Radar-Vision BEVFusion for 3D Object Detection
by: Zhao, Haocheng, et al.
Published: (2024)
by: Zhao, Haocheng, et al.
Published: (2024)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
by: Wen, Zimo, et al.
Published: (2026)
by: Wen, Zimo, et al.
Published: (2026)
Pose-Star: Anatomy-Aware Editing for Open-World Fashion Images
by: Dong, Yuran, et al.
Published: (2025)
by: Dong, Yuran, et al.
Published: (2025)
Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
by: Kwon, Youngchae, et al.
Published: (2025)
by: Kwon, Youngchae, et al.
Published: (2025)
FashionComposer: Compositional Fashion Image Generation
by: Ji, Sihui, et al.
Published: (2024)
by: Ji, Sihui, et al.
Published: (2024)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images
by: Naik, Prithviraj Purushottam, et al.
Published: (2024)
by: Naik, Prithviraj Purushottam, et al.
Published: (2024)
Training-Free Consistency Pipeline for Fashion Repose
by: Aghilar, Potito, et al.
Published: (2025)
by: Aghilar, Potito, et al.
Published: (2025)
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
by: Zhang, Weiyu, et al.
Published: (2026)
by: Zhang, Weiyu, et al.
Published: (2026)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Similar Items
-
FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
by: Pang, Kaicheng, et al.
Published: (2025) -
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025) -
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
by: Yuan, Peng, et al.
Published: (2026) -
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
by: Kong, Xianghao, et al.
Published: (2025) -
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)