Saved in:
Bibliographic Details
Main Authors: Han, Changheon, Kang, Yun Seok, Sim, Yuseop, Park, Hyung Wook, Jun, Martin Byung-Guk
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.07879
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908446492721152
author Han, Changheon
Kang, Yun Seok
Sim, Yuseop
Park, Hyung Wook
Jun, Martin Byung-Guk
author_facet Han, Changheon
Kang, Yun Seok
Sim, Yuseop
Park, Hyung Wook
Jun, Martin Byung-Guk
contents Deep learning-based machine listening is broadening the scope of industrial acoustic analysis for applications like anomaly detection and predictive maintenance, thereby improving manufacturing efficiency and reliability. Nevertheless, its reliance on large, task-specific annotated datasets for every new task limits widespread implementation on shop floors. While emerging sound foundation models aim to alleviate data dependency, they are too large and computationally expensive, requiring cloud infrastructure or high-end hardware that is impractical for on-site, real-time deployment. We address this gap with LISTEN (Lightweight Industrial Sound-representable Transformer for Edge Notification), a kilobyte-sized industrial sound foundation model. Using knowledge distillation, LISTEN runs in real-time on low-cost edge devices. On benchmark downstream tasks, it performs nearly identically to its much larger parent model, even when fine-tuned with minimal datasets and training resource. Beyond the model itself, we demonstrate its real-world utility by integrating LISTEN into a complete machine monitoring framework on an edge device with an Industrial Internet of Things (IIoT) sensor and system, validating its performance and generalization capabilities on a live manufacturing shop floor.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07879
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
Han, Changheon
Kang, Yun Seok
Sim, Yuseop
Park, Hyung Wook
Jun, Martin Byung-Guk
Sound
Audio and Speech Processing
Deep learning-based machine listening is broadening the scope of industrial acoustic analysis for applications like anomaly detection and predictive maintenance, thereby improving manufacturing efficiency and reliability. Nevertheless, its reliance on large, task-specific annotated datasets for every new task limits widespread implementation on shop floors. While emerging sound foundation models aim to alleviate data dependency, they are too large and computationally expensive, requiring cloud infrastructure or high-end hardware that is impractical for on-site, real-time deployment. We address this gap with LISTEN (Lightweight Industrial Sound-representable Transformer for Edge Notification), a kilobyte-sized industrial sound foundation model. Using knowledge distillation, LISTEN runs in real-time on low-cost edge devices. On benchmark downstream tasks, it performs nearly identically to its much larger parent model, even when fine-tuned with minimal datasets and training resource. Beyond the model itself, we demonstrate its real-world utility by integrating LISTEN into a complete machine monitoring framework on an edge device with an Industrial Internet of Things (IIoT) sensor and system, validating its performance and generalization capabilities on a live manufacturing shop floor.
title LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.07879