Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hohman, Fred, Kery, Mary Beth, Ren, Donghao, Moritz, Dominik
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909160189198336
author Hohman, Fred
Kery, Mary Beth
Ren, Donghao
Moritz, Dominik
author_facet Hohman, Fred
Kery, Mary Beth
Ren, Donghao
Moritz, Dominik
contents On-device machine learning (ML) promises to improve the privacy, responsiveness, and proliferation of new, intelligent user experiences by moving ML computation onto everyday personal devices. However, today's large ML models must be drastically compressed to run efficiently on-device, a hurtle that requires deep, yet currently niche expertise. To engage the broader human-centered ML community in on-device ML experiences, we present the results from an interview study with 30 experts at Apple that specialize in producing efficient models. We compile tacit knowledge that experts have developed through practical experience with model compression across different hardware platforms. Our findings offer pragmatic considerations missing from prior work, covering the design process, trade-offs, and technical strategies that go into creating efficient models. Finally, we distill design recommendations for tooling to help ease the difficulty of this work and bring on-device ML into to more widespread practice.
format Preprint
id arxiv_https___arxiv_org_abs_2310_04621
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
Hohman, Fred
Kery, Mary Beth
Ren, Donghao
Moritz, Dominik
Human-Computer Interaction
Artificial Intelligence
Machine Learning
On-device machine learning (ML) promises to improve the privacy, responsiveness, and proliferation of new, intelligent user experiences by moving ML computation onto everyday personal devices. However, today's large ML models must be drastically compressed to run efficiently on-device, a hurtle that requires deep, yet currently niche expertise. To engage the broader human-centered ML community in on-device ML experiences, we present the results from an interview study with 30 experts at Apple that specialize in producing efficient models. We compile tacit knowledge that experts have developed through practical experience with model compression across different hardware platforms. Our findings offer pragmatic considerations missing from prior work, covering the design process, trade-offs, and technical strategies that go into creating efficient models. Finally, we distill design recommendations for tooling to help ease the difficulty of this work and bring on-device ML into to more widespread practice.
title Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
topic Human-Computer Interaction
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2310.04621