Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917629887774720 |
|---|---|
| author | Hohman, Fred Wang, Chaoqun Lee, Jinmook Görtler, Jochen Moritz, Dominik Bigham, Jeffrey P Ren, Zhile Foret, Cecile Shan, Qi Zhang, Xiaoyi |
| author_facet | Hohman, Fred Wang, Chaoqun Lee, Jinmook Görtler, Jochen Moritz, Dominik Bigham, Jeffrey P Ren, Zhile Foret, Cecile Shan, Qi Zhang, Xiaoyi |
| contents | On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria: a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_03085 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference Hohman, Fred Wang, Chaoqun Lee, Jinmook Görtler, Jochen Moritz, Dominik Bigham, Jeffrey P Ren, Zhile Foret, Cecile Shan, Qi Zhang, Xiaoyi Human-Computer Interaction Artificial Intelligence Machine Learning On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria: a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria. |
| title | Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference |
| topic | Human-Computer Interaction Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2404.03085 |