DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911305839935488 |
|---|---|
| author | Kwon, Sangwoo Seo, Seong Hoon Lee, Jae W. Park, Yeonhong |
| author_facet | Kwon, Sangwoo Seo, Seong Hoon Lee, Jae W. Park, Yeonhong |
| contents | How can we effectively handle queries for on-device large language models (LLMs) with varying runtime constraints, such as latency and accuracy? Multi-scale quantization addresses this challenge by enabling memory-efficient runtime model adaptation of LLMs through the overlaying of multiple model variants quantized to different bitwidths. Meanwhile, an important question still remains open-ended: how can models be properly configured to match a target precision or latency? While mixed-precision offers a promising solution, we take this further by leveraging the key observation that the sensitivity of each layer dynamically changes across decoding steps. Building on this insight, we introduce DP-LLM, a novel mechanism that dynamically assigns precision to each layer based on input values. Experimental results across multiple models and benchmarks demonstrate that DP-LLM achieves a superior performance-latency trade-off, outperforming prior approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06041 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment Kwon, Sangwoo Seo, Seong Hoon Lee, Jae W. Park, Yeonhong Machine Learning Artificial Intelligence How can we effectively handle queries for on-device large language models (LLMs) with varying runtime constraints, such as latency and accuracy? Multi-scale quantization addresses this challenge by enabling memory-efficient runtime model adaptation of LLMs through the overlaying of multiple model variants quantized to different bitwidths. Meanwhile, an important question still remains open-ended: how can models be properly configured to match a target precision or latency? While mixed-precision offers a promising solution, we take this further by leveraging the key observation that the sensitivity of each layer dynamically changes across decoding steps. Building on this insight, we introduce DP-LLM, a novel mechanism that dynamically assigns precision to each layer based on input values. Experimental results across multiple models and benchmarks demonstrate that DP-LLM achieves a superior performance-latency trade-off, outperforming prior approaches. |
| title | DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2508.06041 |