LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
Fuente:
arXiv
Saved in:
| Main Authors: | Girella, Federico, Talon, Davide, Liu, Ziyue, Ruan, Zanxi, Wang, Yiming, Cristani, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
by: Liu, Ziyue, et al.
Published: (2025)
by: Liu, Ziyue, et al.
Published: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025)
by: Talon, Davide, et al.
Published: (2025)
Is SAM3 ready for pathology segmentation?
by: Kong, Qiuyu, et al.
Published: (2026)
by: Kong, Qiuyu, et al.
Published: (2026)
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
by: Ruan, Zanxi, et al.
Published: (2026)
by: Ruan, Zanxi, et al.
Published: (2026)
Leveraging Latent Diffusion Models for Training-Free In-Distribution Data Augmentation for Surface Defect Detection
by: Girella, Federico, et al.
Published: (2024)
by: Girella, Federico, et al.
Published: (2024)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
by: Aqeel, Muhammad, et al.
Published: (2025)
by: Aqeel, Muhammad, et al.
Published: (2025)
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
by: Cheng, Wanyin, et al.
Published: (2025)
by: Cheng, Wanyin, et al.
Published: (2025)
ViSketch-GPT: Collaborative Multi-Scale Feature Extraction for Sketch Recognition and Generation
by: Federico, Giulio, et al.
Published: (2025)
by: Federico, Giulio, et al.
Published: (2025)
Multi-Camera Industrial Open-Set Person Re-Identification and Tracking
by: Cunico, Federico, et al.
Published: (2024)
by: Cunico, Federico, et al.
Published: (2024)
Diffusion-based Image Generation for In-distribution Data Augmentation in Surface Defect Detection
by: Capogrosso, Luigi, et al.
Published: (2024)
by: Capogrosso, Luigi, et al.
Published: (2024)
Sketch2NeRF: Multi-view Sketch-guided Text-to-3D Generation
by: Chen, Minglin, et al.
Published: (2024)
by: Chen, Minglin, et al.
Published: (2024)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends
by: Skenderi, Geri, et al.
Published: (2021)
by: Skenderi, Geri, et al.
Published: (2021)
Content-Conditioned Generation of Stylized Free hand Sketches
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
by: Zheng, Zirui, et al.
Published: (2025)
by: Zheng, Zirui, et al.
Published: (2025)
MDiFF: Exploiting Multimodal Score-based Diffusion Models for New Fashion Product Performance Forecasting
by: Avogaro, Andrea, et al.
Published: (2024)
by: Avogaro, Andrea, et al.
Published: (2024)
FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
by: Shi, Chuancheng, et al.
Published: (2025)
by: Shi, Chuancheng, et al.
Published: (2025)
Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation
by: Zorzi, Edoardo, et al.
Published: (2026)
by: Zorzi, Edoardo, et al.
Published: (2026)
Text to Sketch Generation with Multi-Styles
by: Li, Tengjie, et al.
Published: (2025)
by: Li, Tengjie, et al.
Published: (2025)
Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
MCGM: Mask Conditional Text-to-Image Generative Model
by: Skaik, Rami, et al.
Published: (2024)
by: Skaik, Rami, et al.
Published: (2024)
Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
by: Avogaro, Andrea, et al.
Published: (2024)
by: Avogaro, Andrea, et al.
Published: (2024)
New Fashion Products Performance Forecasting: A Survey on Evolutions, Models and Emerging Trends
by: Avogaro, Andrea, et al.
Published: (2025)
by: Avogaro, Andrea, et al.
Published: (2025)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
by: Brioschi, Riccardo, et al.
Published: (2025)
by: Brioschi, Riccardo, et al.
Published: (2025)
SketchTriplet: Self-Supervised Scenarized Sketch-Text-Image Triplet Generation
by: Wu, Zhenbei, et al.
Published: (2024)
by: Wu, Zhenbei, et al.
Published: (2024)
Contrastive Learning-based Multi Modal Architecture for Emoticon Prediction by Employing Image-Text Pairs
by: Pandey, Ananya, et al.
Published: (2024)
by: Pandey, Ananya, et al.
Published: (2024)
AirSketch: Generative Motion to Sketch
by: Lim, Hui Xian Grace, et al.
Published: (2024)
by: Lim, Hui Xian Grace, et al.
Published: (2024)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
TexControl: Sketch-Based Two-Stage Fashion Image Generation Using Diffusion Model
by: Zhang, Yongming, et al.
Published: (2024)
by: Zhang, Yongming, et al.
Published: (2024)
SketchRef: a Multi-Task Evaluation Benchmark for Sketch Synthesis
by: Lin, Xingyue, et al.
Published: (2024)
by: Lin, Xingyue, et al.
Published: (2024)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion
by: Zhang, Ziyue, et al.
Published: (2024)
by: Zhang, Ziyue, et al.
Published: (2024)
Similar Items
-
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
by: Liu, Ziyue, et al.
Published: (2026) -
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
by: Liu, Ziyue, et al.
Published: (2025) -
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025) -
Is SAM3 ready for pathology segmentation?
by: Kong, Qiuyu, et al.
Published: (2026) -
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
by: Ruan, Zanxi, et al.
Published: (2026)