Knowledge boosting during low-latency inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Srinivas, Vidya, Itani, Malek, Chen, Tuochao, Eskimez, Sefik Emre, Yoshioka, Takuya, Gollakota, Shyamnath
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913444542808064
author Srinivas, Vidya
Itani, Malek
Chen, Tuochao
Eskimez, Sefik Emre
Yoshioka, Takuya
Gollakota, Shyamnath
author_facet Srinivas, Vidya
Itani, Malek
Chen, Tuochao
Eskimez, Sefik Emre
Yoshioka, Takuya
Gollakota, Shyamnath
contents Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. Code, dataset, and audio samples available at https://knowledgeboosting.cs.washington.edu/.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11055
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Knowledge boosting during low-latency inference
Srinivas, Vidya
Itani, Malek
Chen, Tuochao
Eskimez, Sefik Emre
Yoshioka, Takuya
Gollakota, Shyamnath
Machine Learning
Sound
Audio and Speech Processing
Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. Code, dataset, and audio samples available at https://knowledgeboosting.cs.washington.edu/.
title Knowledge boosting during low-latency inference
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.11055