A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Agullo, Ferran, Oliveras, Joan, Wang, Chen, Gutierrez-Torre, Alberto, Tardieu, Olivier, Youssef, Alaa, Torres, Jordi, Berral, Josep Ll. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
by: Agullo, Ferran, et al.
Published: (2026)
by: Agullo, Ferran, et al.
Published: (2026)
Towards Pareto Optimal Throughput in Small Language Model Serving
by: Recasens, Pol G., et al.
Published: (2024)
by: Recasens, Pol G., et al.
Published: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)
by: Kumar, Harsh, et al.
Published: (2024)
Profiling Apple Silicon Performance for ML Training
by: Feng, Dahua, et al.
Published: (2025)
by: Feng, Dahua, et al.
Published: (2025)
FRIDA: Free-Rider Detection using Privacy Attacks
by: Recasens, Pol G., et al.
Published: (2024)
by: Recasens, Pol G., et al.
Published: (2024)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
by: Zhang, Quqing, et al.
Published: (2026)
by: Zhang, Quqing, et al.
Published: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
A Model-driven Approach for Continuous Performance Engineering in Microservice-based Systems
by: Cortellessa, Vittorio, et al.
Published: (2023)
by: Cortellessa, Vittorio, et al.
Published: (2023)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
by: Liu, Xiaoxuan, et al.
Published: (2024)
by: Liu, Xiaoxuan, et al.
Published: (2024)
In-Context Bias Propagation in LLM-Based Tabular Data Generation
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
by: Jayakody, Shakya, et al.
Published: (2026)
by: Jayakody, Shakya, et al.
Published: (2026)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026)
by: Beysel, Ege, et al.
Published: (2026)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Beamforming-based Achievable Rate Maximization in ISAC System for Multi-UAV Networking
by: Zhou, Shengcai, et al.
Published: (2025)
by: Zhou, Shengcai, et al.
Published: (2025)
Ecoscape: Fault Tolerance Benchmark for Adaptive Remediation Strategies in Real-Time Edge ML
by: Reiter, Hendrik, et al.
Published: (2025)
by: Reiter, Hendrik, et al.
Published: (2025)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Rethinking Temporal Models for TinyML: LSTM versus 1D-CNN in Resource-Constrained Devices
by: Saha, Bidyut, et al.
Published: (2026)
by: Saha, Bidyut, et al.
Published: (2026)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
by: Sun, Tingyang, et al.
Published: (2026)
by: Sun, Tingyang, et al.
Published: (2026)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
FLEXIS: FLEXible Frequent Subgraph Mining using Maximal Independent Sets
by: Sharma, Akshit, et al.
Published: (2024)
by: Sharma, Akshit, et al.
Published: (2024)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025)
by: Thapa, Krishu K, et al.
Published: (2025)
Fairness in Serving Large Language Models
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
by: Lipshitz, Baraq, et al.
Published: (2025)
by: Lipshitz, Baraq, et al.
Published: (2025)
Machine Learning-driven Autotuning of Graphics Processing Unit Accelerated Computational Fluid Dynamics for Enhanced Performance
by: Xue, Weicheng, et al.
Published: (2023)
by: Xue, Weicheng, et al.
Published: (2023)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
by: Hu, Xiannan, et al.
Published: (2025)
by: Hu, Xiannan, et al.
Published: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
by: Yu, Shan, et al.
Published: (2025)
by: Yu, Shan, et al.
Published: (2025)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
by: Wang, Can, et al.
Published: (2024)
by: Wang, Can, et al.
Published: (2024)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
by: Estevez, Melissa, et al.
Published: (2025)
by: Estevez, Melissa, et al.
Published: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
by: Iglovikov, Vladimir, et al.
Published: (2026)
by: Iglovikov, Vladimir, et al.
Published: (2026)
Similar Items
-
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
by: Agullo, Ferran, et al.
Published: (2026) -
Towards Pareto Optimal Throughput in Small Language Model Serving
by: Recasens, Pol G., et al.
Published: (2024) -
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025) -
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
by: Liu, Qunyou, et al.
Published: (2025) -
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)