AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zicong, Zhang, Kunming, Tang, Guoming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
von: Huang, Hongzhen, et al.
Veröffentlicht: (2025)
von: Huang, Hongzhen, et al.
Veröffentlicht: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
von: Mukherjee, Soutrik, et al.
Veröffentlicht: (2026)
von: Mukherjee, Soutrik, et al.
Veröffentlicht: (2026)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
von: Song, Zhiye, et al.
Veröffentlicht: (2026)
von: Song, Zhiye, et al.
Veröffentlicht: (2026)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
von: Chen, Zixi, et al.
Veröffentlicht: (2025)
von: Chen, Zixi, et al.
Veröffentlicht: (2025)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
WaveTuner: Comprehensive Wavelet Subband Tuning for Time Series Forecasting
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
von: Cui, Yubo, et al.
Veröffentlicht: (2026)
von: Cui, Yubo, et al.
Veröffentlicht: (2026)
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
von: Gopal, Nikhil, et al.
Veröffentlicht: (2026)
von: Gopal, Nikhil, et al.
Veröffentlicht: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
ATFNet: Adaptive Time-Frequency Ensembled Network for Long-term Time Series Forecasting
von: Ye, Hengyu, et al.
Veröffentlicht: (2024)
von: Ye, Hengyu, et al.
Veröffentlicht: (2024)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
von: Roulet, Vincent, et al.
Veröffentlicht: (2024)
von: Roulet, Vincent, et al.
Veröffentlicht: (2024)
Frequency Adaptive Normalization For Non-stationary Time Series Forecasting
von: Ye, Weiwei, et al.
Veröffentlicht: (2024)
von: Ye, Weiwei, et al.
Veröffentlicht: (2024)
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
von: Maliakel, Paul Joe, et al.
Veröffentlicht: (2025)
von: Maliakel, Paul Joe, et al.
Veröffentlicht: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
Performance-driven Constrained Optimal Auto-Tuner for MPC
von: Puigjaner, Albert Gassol, et al.
Veröffentlicht: (2025)
von: Puigjaner, Albert Gassol, et al.
Veröffentlicht: (2025)
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
von: Zhuge, Xiangwen, et al.
Veröffentlicht: (2025)
von: Zhuge, Xiangwen, et al.
Veröffentlicht: (2025)
Sign-Symmetry Learning Rules are Robust Fine-Tuners
von: Berriche, Aymene, et al.
Veröffentlicht: (2025)
von: Berriche, Aymene, et al.
Veröffentlicht: (2025)
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
Accelerating Sparse Transformer Inference on GPU
von: Dai, Wenhao, et al.
Veröffentlicht: (2025)
von: Dai, Wenhao, et al.
Veröffentlicht: (2025)
A Selective Quantization Tuner for ONNX Models
von: Louloudakis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Louloudakis, Nikolaos, et al.
Veröffentlicht: (2025)
Floe: Federated Specialization for Real-Time LLM-SLM Inference
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
von: Gao, Lei, et al.
Veröffentlicht: (2025)
von: Gao, Lei, et al.
Veröffentlicht: (2025)
FusAD: Time-Frequency Fusion with Adaptive Denoising for General Time Series Analysis
von: Zhang, Da, et al.
Veröffentlicht: (2025)
von: Zhang, Da, et al.
Veröffentlicht: (2025)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
Scaling On-Device GPU Inference for Large Generative Models
von: Tang, Jiuqiang, et al.
Veröffentlicht: (2025)
von: Tang, Jiuqiang, et al.
Veröffentlicht: (2025)
Inference-Time Alignment of Diffusion Models with Direct Noise Optimization
von: Tang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Tang, Zhiwei, et al.
Veröffentlicht: (2024)
LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
von: Yi, Jiawei, et al.
Veröffentlicht: (2025)
von: Yi, Jiawei, et al.
Veröffentlicht: (2025)
DemoTuner: Automatic Performance Tuning for Database Management Systems Based on Demonstration Reinforcement Learning
von: Dou, Hui, et al.
Veröffentlicht: (2025)
von: Dou, Hui, et al.
Veröffentlicht: (2025)
A Study of Skews, Imbalances, and Pathological Conditions in LLM Inference Deployment on GPU Clusters detectable from DPU
von: Moye, Javed I. Khan an Henry Uwabor
Veröffentlicht: (2025)
von: Moye, Javed I. Khan an Henry Uwabor
Veröffentlicht: (2025)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
von: Noh, Kangjun, et al.
Veröffentlicht: (2026)
von: Noh, Kangjun, et al.
Veröffentlicht: (2026)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
RAP: Runtime Adaptive Pruning for LLM Inference
von: Liu, Huanrong, et al.
Veröffentlicht: (2025)
von: Liu, Huanrong, et al.
Veröffentlicht: (2025)
ML$^2$Tuner: Efficient Code Tuning via Multi-Level Machine Learning Models
von: Cha, JooHyoung, et al.
Veröffentlicht: (2024)
von: Cha, JooHyoung, et al.
Veröffentlicht: (2024)
FloE: On-the-Fly MoE Inference on Memory-constrained GPU
von: Zhou, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhou, Yuxin, et al.
Veröffentlicht: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
Minimal Batch Adaptive Learning Policy Engine for Real-Time Mid-Price Forecasting in High-Frequency Trading
von: Ntakaris, Adamantios, et al.
Veröffentlicht: (2024)
von: Ntakaris, Adamantios, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
von: Huang, Hongzhen, et al.
Veröffentlicht: (2025) -
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
von: Mukherjee, Soutrik, et al.
Veröffentlicht: (2026) -
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
von: Song, Zhiye, et al.
Veröffentlicht: (2026) -
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
von: Chen, Zixi, et al.
Veröffentlicht: (2025) -
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)