EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Buyuksolak, Oguzhan, Gok, Alican, Okman, Osman Erman
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912842728341504
author Buyuksolak, Oguzhan
Gok, Alican
Okman, Osman Erman
author_facet Buyuksolak, Oguzhan
Gok, Alican
Okman, Osman Erman
contents We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16316
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
Buyuksolak, Oguzhan
Gok, Alican
Okman, Osman Erman
Audio and Speech Processing
Computation and Language
Sound
We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.
title EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2601.16316