LLMs on a Budget? Say HOLA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siddiqui, Zohaib Hasan, Gao, Jiechao, Shabbir, Ebad, Azeez, Mohammad Anas, Ali, Rafiq, Kashyap, Gautam Siddharth, Naseem, Usman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912637249388544
author Siddiqui, Zohaib Hasan
Gao, Jiechao
Shabbir, Ebad
Azeez, Mohammad Anas
Ali, Rafiq
Kashyap, Gautam Siddharth
Naseem, Usman
author_facet Siddiqui, Zohaib Hasan
Gao, Jiechao
Shabbir, Ebad
Azeez, Mohammad Anas
Ali, Rafiq
Kashyap, Gautam Siddharth
Naseem, Usman
contents Running Large Language Models (LLMs) on edge devices is constrained by high compute and memory demands posing a barrier for real-time applications in sectors like healthcare, education, and embedded systems. Current solutions such as quantization, pruning, and retrieval-augmented generation (RAG) offer only partial optimizations and often compromise on speed or accuracy. We introduce HOLA, an end-to-end optimization framework for efficient LLM deployment. Internally, it leverages Hierarchical Speculative Decoding (HSD) for faster inference without quality loss. Externally, AdaComp-RAG adjusts retrieval complexity based on context needs. Together with LoBi, which blends structured pruning (LoRA) and quantization, HOLA delivers significant gains: 17.6% EMA on GSM8K, 10.5% MCA on ARC, and reduced latency and memory on edge devices like Jetson Nano--proving both scalable and production-ready.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs on a Budget? Say HOLA
Siddiqui, Zohaib Hasan
Gao, Jiechao
Shabbir, Ebad
Azeez, Mohammad Anas
Ali, Rafiq
Kashyap, Gautam Siddharth
Naseem, Usman
Machine Learning
Artificial Intelligence
Computation and Language
Running Large Language Models (LLMs) on edge devices is constrained by high compute and memory demands posing a barrier for real-time applications in sectors like healthcare, education, and embedded systems. Current solutions such as quantization, pruning, and retrieval-augmented generation (RAG) offer only partial optimizations and often compromise on speed or accuracy. We introduce HOLA, an end-to-end optimization framework for efficient LLM deployment. Internally, it leverages Hierarchical Speculative Decoding (HSD) for faster inference without quality loss. Externally, AdaComp-RAG adjusts retrieval complexity based on context needs. Together with LoBi, which blends structured pruning (LoRA) and quantization, HOLA delivers significant gains: 17.6% EMA on GSM8K, 10.5% MCA on ARC, and reduced latency and memory on edge devices like Jetson Nano--proving both scalable and production-ready.
title LLMs on a Budget? Say HOLA
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.18952