Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tanaka, Daichi, Karasawa, Takumi, Takenouchi, Shu, Kawakami, Rei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915346517065728
author Tanaka, Daichi
Karasawa, Takumi
Takenouchi, Shu
Kawakami, Rei
author_facet Tanaka, Daichi
Karasawa, Takumi
Takenouchi, Shu
Kawakami, Rei
contents Recycling steel scrap can reduce carbon dioxide (CO2) emissions from the steel industry. However, a significant challenge in steel scrap recycling is the inclusion of impurities other than steel. To address this issue, we propose vision-language-model-based anomaly detection where a model is finetuned in a supervised manner, enabling it to handle niche objects effectively. This model enables automated detection of anomalies at a fine-grained level within steel scrap. Specifically, we finetune the image encoder, equipped with multi-scale mechanism and text prompts aligned with both normal and anomaly images. The finetuning process trains these modules using a multiclass classification as the supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
Tanaka, Daichi
Karasawa, Takumi
Takenouchi, Shu
Kawakami, Rei
Computer Vision and Pattern Recognition
Recycling steel scrap can reduce carbon dioxide (CO2) emissions from the steel industry. However, a significant challenge in steel scrap recycling is the inclusion of impurities other than steel. To address this issue, we propose vision-language-model-based anomaly detection where a model is finetuned in a supervised manner, enabling it to handle niche objects effectively. This model enables automated detection of anomalies at a fine-grained level within steel scrap. Specifically, we finetune the image encoder, equipped with multi-scale mechanism and text prompts aligned with both normal and anomaly images. The finetuning process trains these modules using a multiclass classification as the supervision.
title Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13282