DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
Fuente:
arXiv
Saved in:
| Main Authors: | Kwon, Sangwoo, Seo, Seong Hoon, Lee, Jae W., Park, Yeonhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
by: Seo, Sangwoo, et al.
Published: (2025)
by: Seo, Sangwoo, et al.
Published: (2025)
Interpretable Prototype-based Graph Information Bottleneck
by: Seo, Sangwoo, et al.
Published: (2023)
by: Seo, Sangwoo, et al.
Published: (2023)
Unsupervised Episode Generation for Graph Meta-learning
by: Jung, Jihyeong, et al.
Published: (2023)
by: Jung, Jihyeong, et al.
Published: (2023)
Disentangling Hyperedges through the Lens of Category Theory
by: Lee, Yoonho, et al.
Published: (2025)
by: Lee, Yoonho, et al.
Published: (2025)
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Target Circuit Matching in Large-Scale Netlists using GNN-Based Region Prediction
by: Seo, Sangwoo, et al.
Published: (2025)
by: Seo, Sangwoo, et al.
Published: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Have it your way: Individualized Privacy Assignment for DP-SGD
by: Boenisch, Franziska, et al.
Published: (2023)
by: Boenisch, Franziska, et al.
Published: (2023)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
by: Shen, Yiqun, et al.
Published: (2025)
by: Shen, Yiqun, et al.
Published: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
by: Park, Geon, et al.
Published: (2026)
by: Park, Geon, et al.
Published: (2026)
Uncertainty Calibration with Energy Based Instance-wise Scaling in the Wild Dataset
by: Kim, Mijoo, et al.
Published: (2024)
by: Kim, Mijoo, et al.
Published: (2024)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
RAP: Runtime Adaptive Pruning for LLM Inference
by: Liu, Huanrong, et al.
Published: (2025)
by: Liu, Huanrong, et al.
Published: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)
by: Xia, Haojun, et al.
Published: (2025)
Efficient Process Reward Modeling via Contrastive Mutual Information
by: Lee, Nakyung, et al.
Published: (2026)
by: Lee, Nakyung, et al.
Published: (2026)
A Layer-wise Analysis of Supervised Fine-Tuning
by: Zhao, Qinghua, et al.
Published: (2026)
by: Zhao, Qinghua, et al.
Published: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
Non-linear Interventions on Large Language Models
by: Kim, Sangwoo
Published: (2026)
by: Kim, Sangwoo
Published: (2026)
HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
by: Zhao, Huaqin, et al.
Published: (2024)
by: Zhao, Huaqin, et al.
Published: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
by: Fartale, Harshwardhan, et al.
Published: (2025)
by: Fartale, Harshwardhan, et al.
Published: (2025)
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
by: Gwak, Minju, et al.
Published: (2026)
by: Gwak, Minju, et al.
Published: (2026)
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
Stochastic Layer-wise Learning: Scalable and Efficient Alternative to Backpropagation
by: Yin, Bojian, et al.
Published: (2025)
by: Yin, Bojian, et al.
Published: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Resource-efficient Layer-wise Federated Self-supervised Learning
by: Tun, Ye Lin, et al.
Published: (2024)
by: Tun, Ye Lin, et al.
Published: (2024)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
Efficient Knowledge Deletion from Trained Models through Layer-wise Partial Machine Unlearning
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
by: Park, Jaehyun, et al.
Published: (2025)
by: Park, Jaehyun, et al.
Published: (2025)
Information-Theoretic Greedy Layer-wise Training for Traffic Sign Recognition
by: Lyu, Shuyan, et al.
Published: (2025)
by: Lyu, Shuyan, et al.
Published: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
Similar Items
-
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
by: Seo, Sangwoo, et al.
Published: (2025) -
Interpretable Prototype-based Graph Information Bottleneck
by: Seo, Sangwoo, et al.
Published: (2023) -
Unsupervised Episode Generation for Graph Meta-learning
by: Jung, Jihyeong, et al.
Published: (2023) -
Disentangling Hyperedges through the Lens of Category Theory
by: Lee, Yoonho, et al.
Published: (2025) -
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024)