Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hohman, Fred, Wang, Chaoqun, Lee, Jinmook, Görtler, Jochen, Moritz, Dominik, Bigham, Jeffrey P, Ren, Zhile, Foret, Cecile, Shan, Qi, Zhang, Xiaoyi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917629887774720
author Hohman, Fred
Wang, Chaoqun
Lee, Jinmook
Görtler, Jochen
Moritz, Dominik
Bigham, Jeffrey P
Ren, Zhile
Foret, Cecile
Shan, Qi
Zhang, Xiaoyi
author_facet Hohman, Fred
Wang, Chaoqun
Lee, Jinmook
Görtler, Jochen
Moritz, Dominik
Bigham, Jeffrey P
Ren, Zhile
Foret, Cecile
Shan, Qi
Zhang, Xiaoyi
contents On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria: a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03085
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
Hohman, Fred
Wang, Chaoqun
Lee, Jinmook
Görtler, Jochen
Moritz, Dominik
Bigham, Jeffrey P
Ren, Zhile
Foret, Cecile
Shan, Qi
Zhang, Xiaoyi
Human-Computer Interaction
Artificial Intelligence
Machine Learning
On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria: a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria.
title Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
topic Human-Computer Interaction
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.03085