Salvato in:
Dettagli Bibliografici
Autore principale: Banasik, Spencer
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2506.01827
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908389784682496
author Banasik, Spencer
author_facet Banasik, Spencer
contents As machine learning algorithms are shown to be an increasingly valuable tool, the demand for their access has grown accordingly. Oftentimes, it is infeasible to run inference with larger models without an accelerator, which may be unavailable in environments that have constraints such as energy consumption, security, or cost. To increase the availability of these models, we aim to improve the LLM inference speed on a CPU-only environment by modifying the cache architecture. To determine what improvements could be made, we conducted two experiments using Llama.cpp and the QWEN model: running various cache configurations and evaluating their performance, and outputting a trace of the memory footprint. Using these experiments, we investigate the memory access patterns and performance characteristics to identify potential optimizations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
Banasik, Spencer
Machine Learning
Hardware Architecture
As machine learning algorithms are shown to be an increasingly valuable tool, the demand for their access has grown accordingly. Oftentimes, it is infeasible to run inference with larger models without an accelerator, which may be unavailable in environments that have constraints such as energy consumption, security, or cost. To increase the availability of these models, we aim to improve the LLM inference speed on a CPU-only environment by modifying the cache architecture. To determine what improvements could be made, we conducted two experiments using Llama.cpp and the QWEN model: running various cache configurations and evaluating their performance, and outputting a trace of the memory footprint. Using these experiments, we investigate the memory access patterns and performance characteristics to identify potential optimizations.
title Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2506.01827