LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singh, Shubhr, Benetos, Emmanouil, Phan, Huy, Stowell, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913669465505792
author Singh, Shubhr
Benetos, Emmanouil
Phan, Huy
Stowell, Dan
author_facet Singh, Shubhr
Benetos, Emmanouil
Phan, Huy
Stowell, Dan
contents Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to process the higher-order relations essential for identifying distinct audio objects. To address this limitation, this work introduces the Local- Higher Order Graph Neural Network (LHGNN), a graph based model that enhances feature understanding by integrating local neighbourhood information with higher-order data from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audio relationships. Evaluation of the model on three publicly available audio datasets shows that it outperforms Transformer-based models across all benchmarks while operating with substantially fewer parameters. Moreover, LHGNN demonstrates a distinct advantage in scenarios lacking ImageNet pretraining, establishing its effectiveness and efficiency in environments where extensive pretraining data is unavailable.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03464
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
Singh, Shubhr
Benetos, Emmanouil
Phan, Huy
Stowell, Dan
Sound
Artificial Intelligence
Audio and Speech Processing
Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to process the higher-order relations essential for identifying distinct audio objects. To address this limitation, this work introduces the Local- Higher Order Graph Neural Network (LHGNN), a graph based model that enhances feature understanding by integrating local neighbourhood information with higher-order data from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audio relationships. Evaluation of the model on three publicly available audio datasets shows that it outperforms Transformer-based models across all benchmarks while operating with substantially fewer parameters. Moreover, LHGNN demonstrates a distinct advantage in scenarios lacking ImageNet pretraining, establishing its effectiveness and efficiency in environments where extensive pretraining data is unavailable.
title LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2501.03464