EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912842728341504 |
|---|---|
| author | Buyuksolak, Oguzhan Gok, Alican Okman, Osman Erman |
| author_facet | Buyuksolak, Oguzhan Gok, Alican Okman, Osman Erman |
| contents | We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_16316 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting Buyuksolak, Oguzhan Gok, Alican Okman, Osman Erman Audio and Speech Processing Computation and Language Sound We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters. |
| title | EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2601.16316 |