Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mitsis, Stavros, Hadjikyriakos, Ermos, Ibrahim, Humaid, Neofytou, Savvas, Raman, Shashwat, Myles, James, Kanjo, Eiman
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911222884990976
author Mitsis, Stavros
Hadjikyriakos, Ermos
Ibrahim, Humaid
Neofytou, Savvas
Raman, Shashwat
Myles, James
Kanjo, Eiman
author_facet Mitsis, Stavros
Hadjikyriakos, Ermos
Ibrahim, Humaid
Neofytou, Savvas
Raman, Shashwat
Myles, James
Kanjo, Eiman
contents Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict de-escalation, and responsive wearables, where cloud-based solutions are impractical. Multimodal emotion recognition has advanced through deep learning, but most systems remain unsuitable for deployment on ultra-constrained edge devices. Prior work typically relies on powerful hardware, lacks real-time performance, or uses unimodal input. This paper addresses that gap by presenting a hardware-aware emotion recognition system that combines acoustic and linguistic features using a late-fusion architecture optimised for Edge TPU. The design integrates a quantised transformer-based acoustic model with frozen keyword embeddings from a DSResNet-SE network, enabling real-time inference within a 1.8MB memory budget and 21-23ms latency. The pipeline ensures spectrogram alignment between training and deployment using MicroFrontend and MLTK. Evaluation on re-recorded, segmented IEMOCAP samples captured through the Coral Dev Board Micro microphone shows a 6.3% macro F1 improvement over unimodal baselines. This work demonstrates that accurate, real-time multimodal emotion inference is achievable on microcontroller-class edge platforms through task-specific fusion and hardware-guided model design.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
Mitsis, Stavros
Hadjikyriakos, Ermos
Ibrahim, Humaid
Neofytou, Savvas
Raman, Shashwat
Myles, James
Kanjo, Eiman
Sound
Machine Learning
Audio and Speech Processing
Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict de-escalation, and responsive wearables, where cloud-based solutions are impractical. Multimodal emotion recognition has advanced through deep learning, but most systems remain unsuitable for deployment on ultra-constrained edge devices. Prior work typically relies on powerful hardware, lacks real-time performance, or uses unimodal input. This paper addresses that gap by presenting a hardware-aware emotion recognition system that combines acoustic and linguistic features using a late-fusion architecture optimised for Edge TPU. The design integrates a quantised transformer-based acoustic model with frozen keyword embeddings from a DSResNet-SE network, enabling real-time inference within a 1.8MB memory budget and 21-23ms latency. The pipeline ensures spectrogram alignment between training and deployment using MicroFrontend and MLTK. Evaluation on re-recorded, segmented IEMOCAP samples captured through the Coral Dev Board Micro microphone shows a 6.3% macro F1 improvement over unimodal baselines. This work demonstrates that accurate, real-time multimodal emotion inference is achievable on microcontroller-class edge platforms through task-specific fusion and hardware-guided model design.
title Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2510.18036