MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Liman, Zhong, Hanyang, Wang, Tianyuan, Luo, Shan, Zhu, Jihong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915545306103808
author Wang, Liman
Zhong, Hanyang
Wang, Tianyuan
Luo, Shan
Zhu, Jihong
author_facet Wang, Liman
Zhong, Hanyang
Wang, Tianyuan
Luo, Shan
Zhu, Jihong
contents Choosing appropriate fabrics is critical for meeting functional and quality demands in robotic textile manufacturing, apparel production, and smart retail. We propose MLLM-Fabric, a robotic framework leveraging multimodal large language models (MLLMs) for fabric sorting and selection. Built on a multimodal robotic platform, the system is trained through supervised fine-tuning and explanation-guided distillation to rank fabric properties. We also release a dataset of 220 diverse fabrics, each with RGB images and synchronized visuotactile and pressure data. Experiments show that our Fabric-Llama-90B consistently outperforms pretrained vision-language baselines in both attribute ranking and selection reliability. Code and dataset are publicly available at https://github.com/limanwang/MLLM-Fabric.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
Wang, Liman
Zhong, Hanyang
Wang, Tianyuan
Luo, Shan
Zhu, Jihong
Robotics
Artificial Intelligence
Choosing appropriate fabrics is critical for meeting functional and quality demands in robotic textile manufacturing, apparel production, and smart retail. We propose MLLM-Fabric, a robotic framework leveraging multimodal large language models (MLLMs) for fabric sorting and selection. Built on a multimodal robotic platform, the system is trained through supervised fine-tuning and explanation-guided distillation to rank fabric properties. We also release a dataset of 220 diverse fabrics, each with RGB images and synchronized visuotactile and pressure data. Experiments show that our Fabric-Llama-90B consistently outperforms pretrained vision-language baselines in both attribute ranking and selection reliability. Code and dataset are publicly available at https://github.com/limanwang/MLLM-Fabric.
title MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2507.04351