Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Ziyue, Talon, Davide, Girella, Federico, Ruan, Zanxi, Mondo, Mattia, Bazzani, Loris, Wang, Yiming, Cristani, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
von: Girella, Federico, et al.
Veröffentlicht: (2025)
von: Girella, Federico, et al.
Veröffentlicht: (2025)
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
von: Talon, Davide, et al.
Veröffentlicht: (2025)
von: Talon, Davide, et al.
Veröffentlicht: (2025)
Is SAM3 ready for pathology segmentation?
von: Kong, Qiuyu, et al.
Veröffentlicht: (2026)
von: Kong, Qiuyu, et al.
Veröffentlicht: (2026)
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
von: Ruan, Zanxi, et al.
Veröffentlicht: (2026)
von: Ruan, Zanxi, et al.
Veröffentlicht: (2026)
Leveraging Latent Diffusion Models for Training-Free In-Distribution Data Augmentation for Surface Defect Detection
von: Girella, Federico, et al.
Veröffentlicht: (2024)
von: Girella, Federico, et al.
Veröffentlicht: (2024)
Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation
von: Zorzi, Edoardo, et al.
Veröffentlicht: (2026)
von: Zorzi, Edoardo, et al.
Veröffentlicht: (2026)
Diffusion-based Image Generation for In-distribution Data Augmentation in Surface Defect Detection
von: Capogrosso, Luigi, et al.
Veröffentlicht: (2024)
von: Capogrosso, Luigi, et al.
Veröffentlicht: (2024)
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
von: Cheng, Wanyin, et al.
Veröffentlicht: (2025)
von: Cheng, Wanyin, et al.
Veröffentlicht: (2025)
Multi-Camera Industrial Open-Set Person Re-Identification and Tracking
von: Cunico, Federico, et al.
Veröffentlicht: (2024)
von: Cunico, Federico, et al.
Veröffentlicht: (2024)
Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends
von: Skenderi, Geri, et al.
Veröffentlicht: (2021)
von: Skenderi, Geri, et al.
Veröffentlicht: (2021)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
von: Baldrati, Alberto, et al.
Veröffentlicht: (2024)
von: Baldrati, Alberto, et al.
Veröffentlicht: (2024)
Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2025)
Interactive Episodic Memory with User Feedback
von: Subedi, Nikesh, et al.
Veröffentlicht: (2026)
von: Subedi, Nikesh, et al.
Veröffentlicht: (2026)
MDiFF: Exploiting Multimodal Score-based Diffusion Models for New Fashion Product Performance Forecasting
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
von: Das, Deepayan, et al.
Veröffentlicht: (2024)
Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
New Fashion Products Performance Forecasting: A Survey on Evolutions, Models and Emerging Trends
von: Avogaro, Andrea, et al.
Veröffentlicht: (2025)
von: Avogaro, Andrea, et al.
Veröffentlicht: (2025)
UniCoRN: Unified Commented Retrieval Network with LMMs
von: Jaritz, Maximilian, et al.
Veröffentlicht: (2025)
von: Jaritz, Maximilian, et al.
Veröffentlicht: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
von: Shi, Chuancheng, et al.
Veröffentlicht: (2025)
von: Shi, Chuancheng, et al.
Veröffentlicht: (2025)
Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval
von: Wang, Ziwei, et al.
Veröffentlicht: (2024)
von: Wang, Ziwei, et al.
Veröffentlicht: (2024)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
Text to Sketch Generation with Multi-Styles
von: Li, Tengjie, et al.
Veröffentlicht: (2025)
von: Li, Tengjie, et al.
Veröffentlicht: (2025)
ViSketch-GPT: Collaborative Multi-Scale Feature Extraction for Sketch Recognition and Generation
von: Federico, Giulio, et al.
Veröffentlicht: (2025)
von: Federico, Giulio, et al.
Veröffentlicht: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
TexControl: Sketch-Based Two-Stage Fashion Image Generation Using Diffusion Model
von: Zhang, Yongming, et al.
Veröffentlicht: (2024)
von: Zhang, Yongming, et al.
Veröffentlicht: (2024)
SketchTriplet: Self-Supervised Scenarized Sketch-Text-Image Triplet Generation
von: Wu, Zhenbei, et al.
Veröffentlicht: (2024)
von: Wu, Zhenbei, et al.
Veröffentlicht: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
von: Das, Deepayan, et al.
Veröffentlicht: (2025)
FashionComposer: Compositional Fashion Image Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
Sketch2NeRF: Multi-view Sketch-guided Text-to-3D Generation
von: Chen, Minglin, et al.
Veröffentlicht: (2024)
von: Chen, Minglin, et al.
Veröffentlicht: (2024)
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
von: Lin, Haichuan, et al.
Veröffentlicht: (2025)
von: Lin, Haichuan, et al.
Veröffentlicht: (2025)
ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
Let Human Sketches Help: Empowering Challenging Image Segmentation Task with Freehand Sketches
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models
von: Qin, Qice, et al.
Veröffentlicht: (2024)
von: Qin, Qice, et al.
Veröffentlicht: (2024)
Learning based Ge'ez character handwritten recognition
von: Yimer, Hailemicael Lulseged, et al.
Veröffentlicht: (2024)
von: Yimer, Hailemicael Lulseged, et al.
Veröffentlicht: (2024)
How to Take a Memorable Picture? Empowering Users with Actionable Feedback
von: Laiti, Francesco, et al.
Veröffentlicht: (2026)
von: Laiti, Francesco, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
von: Girella, Federico, et al.
Veröffentlicht: (2025) -
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025) -
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
von: Talon, Davide, et al.
Veröffentlicht: (2025) -
Is SAM3 ready for pathology segmentation?
von: Kong, Qiuyu, et al.
Veröffentlicht: (2026) -
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
von: Ruan, Zanxi, et al.
Veröffentlicht: (2026)