TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Yujie, Zhang, Xuanyu, Li, Weiqi, Zhang, Jian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913932243894272
author Hu, Yujie
Zhang, Xuanyu
Li, Weiqi
Zhang, Jian
author_facet Hu, Yujie
Zhang, Xuanyu
Li, Weiqi
Zhang, Jian
contents Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily relied on end-to-end networks to perform single try-on tasks, lacking versatility and flexibility. We propose TalkFashion, an intelligent try-on assistant that leverages the powerful comprehension capabilities of large language models to analyze user instructions and determine which task to execute, thereby activating different processing pipelines accordingly. Additionally, we introduce an instruction-based local repainting model that eliminates the need for users to manually provide masks. With the help of multi-modal models, this approach achieves fully automated local editings, enhancing the flexibility of editing tasks. The experimental results demonstrate better semantic consistency and visual quality compared to the current methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
Hu, Yujie
Zhang, Xuanyu
Li, Weiqi
Zhang, Jian
Computer Vision and Pattern Recognition
Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily relied on end-to-end networks to perform single try-on tasks, lacking versatility and flexibility. We propose TalkFashion, an intelligent try-on assistant that leverages the powerful comprehension capabilities of large language models to analyze user instructions and determine which task to execute, thereby activating different processing pipelines accordingly. Additionally, we introduce an instruction-based local repainting model that eliminates the need for users to manually provide masks. With the help of multi-modal models, this approach achieves fully automated local editings, enhancing the flexibility of editing tasks. The experimental results demonstrate better semantic consistency and visual quality compared to the current methods.
title TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05790