ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gu, Yile, Kadekodi, Rohan, Nguyen, Hoang, Kamahori, Keisuke, Liu, Yiyu, Kasikci, Baris
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915354570129408
author Gu, Yile
Kadekodi, Rohan
Nguyen, Hoang
Kamahori, Keisuke
Liu, Yiyu
Kasikci, Baris
author_facet Gu, Yile
Kadekodi, Rohan
Nguyen, Hoang
Kamahori, Keisuke
Liu, Yiyu
Kasikci, Baris
contents The recent shift in Generative AI (GenAI) applications from cloud-only environments to end-user devices introduces new challenges in resource management, system efficiency, and user experience. This paper presents ConsumerBench, a comprehensive benchmarking framework designed to evaluate the system efficiency and response time of GenAI models running on end-user devices. Unlike existing benchmarks that assume exclusive model access on dedicated GPUs, ConsumerBench simulates realistic multi-application scenarios executing concurrently on constrained hardware. Furthermore, ConsumerBench supports customizable workflows that simulate complex tasks requiring coordination among multiple applications. ConsumerBench captures both application-level metrics, including latency and Service Level Objective (SLO) attainment, and system-level metrics like CPU/GPU utilization and memory bandwidth. Through extensive experiments, ConsumerBench reveals inefficiencies in resource sharing, unfair scheduling under greedy allocation, and performance pitfalls of static model server configurations. The paper also provides practical insights for model developers and system designers, highlighting the benefits of custom kernels tailored to consumer-grade GPU architectures and the value of implementing SLO-aware scheduling strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
Gu, Yile
Kadekodi, Rohan
Nguyen, Hoang
Kamahori, Keisuke
Liu, Yiyu
Kasikci, Baris
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Operating Systems
The recent shift in Generative AI (GenAI) applications from cloud-only environments to end-user devices introduces new challenges in resource management, system efficiency, and user experience. This paper presents ConsumerBench, a comprehensive benchmarking framework designed to evaluate the system efficiency and response time of GenAI models running on end-user devices. Unlike existing benchmarks that assume exclusive model access on dedicated GPUs, ConsumerBench simulates realistic multi-application scenarios executing concurrently on constrained hardware. Furthermore, ConsumerBench supports customizable workflows that simulate complex tasks requiring coordination among multiple applications. ConsumerBench captures both application-level metrics, including latency and Service Level Objective (SLO) attainment, and system-level metrics like CPU/GPU utilization and memory bandwidth. Through extensive experiments, ConsumerBench reveals inefficiencies in resource sharing, unfair scheduling under greedy allocation, and performance pitfalls of static model server configurations. The paper also provides practical insights for model developers and system designers, highlighting the benefits of custom kernels tailored to consumer-grade GPU architectures and the value of implementing SLO-aware scheduling strategies.
title ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Operating Systems
url https://arxiv.org/abs/2506.17538