Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grau-Haro, Jordi, Ribes-Serrano, Ruben, Naranjo-Alcazar, Javier, Garcia-Ballesteros, Marta, Zuccarello, Pedro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918144034996224
author Grau-Haro, Jordi
Ribes-Serrano, Ruben
Naranjo-Alcazar, Javier
Garcia-Ballesteros, Marta
Zuccarello, Pedro
author_facet Grau-Haro, Jordi
Ribes-Serrano, Ruben
Naranjo-Alcazar, Javier
Garcia-Ballesteros, Marta
Zuccarello, Pedro
contents Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to computational efficiency and thermal management. In this paper, a comprehensive evaluation of multiple convolutional neural network (CNN) architectures for audio tagging on the Raspberry Pi is conducted, encompassing all 1D and 2D models from the Pretrained Audio Neural Networks (PANNs) framework, a ConvNeXt-based model adapted for audio classification, as well as MobileNetV3 architectures. In addition, two PANNs-derived networks, CNN9 and CNN13, recently proposed, are also evaluated. To enhance deployment efficiency and portability across diverse hardware platforms, all models are converted to the Open Neural Network Exchange (ONNX) format. Unlike previous works that focus on a single model, our analysis encompasses a broader range of architectures and involves continuous 24-hour inference sessions to assess performance stability. Our experiments reveal that, with appropriate model selection and optimization, it is possible to maintain consistent inference latency and manage thermal behavior effectively over extended periods. These findings provide valuable insights for deploying audio tagging models in real-world edge computing scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
Grau-Haro, Jordi
Ribes-Serrano, Ruben
Naranjo-Alcazar, Javier
Garcia-Ballesteros, Marta
Zuccarello, Pedro
Sound
Artificial Intelligence
Audio and Speech Processing
Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to computational efficiency and thermal management. In this paper, a comprehensive evaluation of multiple convolutional neural network (CNN) architectures for audio tagging on the Raspberry Pi is conducted, encompassing all 1D and 2D models from the Pretrained Audio Neural Networks (PANNs) framework, a ConvNeXt-based model adapted for audio classification, as well as MobileNetV3 architectures. In addition, two PANNs-derived networks, CNN9 and CNN13, recently proposed, are also evaluated. To enhance deployment efficiency and portability across diverse hardware platforms, all models are converted to the Open Neural Network Exchange (ONNX) format. Unlike previous works that focus on a single model, our analysis encompasses a broader range of architectures and involves continuous 24-hour inference sessions to assess performance stability. Our experiments reveal that, with appropriate model selection and optimization, it is possible to maintain consistent inference latency and manage thermal behavior effectively over extended periods. These findings provide valuable insights for deploying audio tagging models in real-world edge computing scenarios.
title Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.14049