Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ni, TsaiChing, Chen, ZhenQi, Yang, YuanFu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915719602503680
author Ni, TsaiChing
Chen, ZhenQi
Yang, YuanFu
author_facet Ni, TsaiChing
Chen, ZhenQi
Yang, YuanFu
contents We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution real-world defects spanning over 60 material categories and more than 400 defect types, each accompanied by expert-verified annotations and fine-grained textual descriptions detailing defect location, severity, and contextual attributes. This dataset enables a wide spectrum of applications, including classification, segmentation, retrieval, captioning, and generative modeling. Building upon IMDD-1M, we train a diffusion-based vision-language foundation model from scratch, specifically tailored for industrial scenarios. The model serves as a generalizable foundation that can be efficiently adapted to specialized domains through lightweight fine-tuning. With less than 5% of the task-specific data required by dedicated expert models, it achieves comparable performance, highlighting the potential of data-efficient foundation model adaptation for industrial inspection and generation, paving the way for scalable, domain-adaptive, and knowledge-grounded manufacturing intelligence. Additional details and resources can be found in this URL: https://ninaneon.github.io/projectpage/
format Preprint
id arxiv_https___arxiv_org_abs_2512_24160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
Ni, TsaiChing
Chen, ZhenQi
Yang, YuanFu
Computer Vision and Pattern Recognition
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution real-world defects spanning over 60 material categories and more than 400 defect types, each accompanied by expert-verified annotations and fine-grained textual descriptions detailing defect location, severity, and contextual attributes. This dataset enables a wide spectrum of applications, including classification, segmentation, retrieval, captioning, and generative modeling. Building upon IMDD-1M, we train a diffusion-based vision-language foundation model from scratch, specifically tailored for industrial scenarios. The model serves as a generalizable foundation that can be efficiently adapted to specialized domains through lightweight fine-tuning. With less than 5% of the task-specific data required by dedicated expert models, it achieves comparable performance, highlighting the potential of data-efficient foundation model adaptation for industrial inspection and generation, paving the way for scalable, domain-adaptive, and knowledge-grounded manufacturing intelligence. Additional details and resources can be found in this URL: https://ninaneon.github.io/projectpage/
title Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.24160