LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Crumpacker, Katelyn, Nikolopoulos, Dimitrios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
by: Dhulipala, Hridya, et al.
Published: (2025)
by: Dhulipala, Hridya, et al.
Published: (2025)
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
by: Alizadeh, Negar, et al.
Published: (2024)
by: Alizadeh, Negar, et al.
Published: (2024)
Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters
by: Zine, Nada, et al.
Published: (2026)
by: Zine, Nada, et al.
Published: (2026)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
Is Hyper-Parameter Optimization Different for Software Analytics?
by: Yedida, Rahul, et al.
Published: (2024)
by: Yedida, Rahul, et al.
Published: (2024)
Themisto: Jupyter-Based Runtime Benchmark
by: Grotov, Konstantin, et al.
Published: (2025)
by: Grotov, Konstantin, et al.
Published: (2025)
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
by: Galliamov, Karim, et al.
Published: (2024)
by: Galliamov, Karim, et al.
Published: (2024)
Shapley-Guided Neural Repair Approach via Derivative-Free Optimization
by: Sun, Xinyu, et al.
Published: (2026)
by: Sun, Xinyu, et al.
Published: (2026)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
by: Su, Jianhao, et al.
Published: (2026)
by: Su, Jianhao, et al.
Published: (2026)
MARCO: Multi-Agent Code Optimization with Real-Time Knowledge Integration for High-Performance Computing
by: Rahman, Asif, et al.
Published: (2025)
by: Rahman, Asif, et al.
Published: (2025)
Runtime Anomaly Detection for Drones: An Integrated Rule-Mining and Unsupervised-Learning Approach
by: Tan, Ivan, et al.
Published: (2025)
by: Tan, Ivan, et al.
Published: (2025)
PRIMG : Efficient LLM-driven Test Generation Using Mutant Prioritization
by: Bouafif, Mohamed Salah, et al.
Published: (2025)
by: Bouafif, Mohamed Salah, et al.
Published: (2025)
Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
by: Zeng, Wenhao, et al.
Published: (2025)
by: Zeng, Wenhao, et al.
Published: (2025)
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
by: Ryan, Gabriel, et al.
Published: (2024)
by: Ryan, Gabriel, et al.
Published: (2024)
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
by: Tsai, Yun-Da, et al.
Published: (2024)
by: Tsai, Yun-Da, et al.
Published: (2024)
Leveraging Reward Models for Guiding Code Review Comment Generation
by: Sghaier, Oussama Ben, et al.
Published: (2025)
by: Sghaier, Oussama Ben, et al.
Published: (2025)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
by: Hu, Boyue Caroline, et al.
Published: (2025)
by: Hu, Boyue Caroline, et al.
Published: (2025)
An Efficient Model Maintenance Approach for MLOps
by: Majidi, Forough, et al.
Published: (2024)
by: Majidi, Forough, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
by: Storhaug, André, et al.
Published: (2024)
by: Storhaug, André, et al.
Published: (2024)
iServe: An Intent-based Serving System for LLMs
by: Liakopoulos, Dimitrios, et al.
Published: (2025)
by: Liakopoulos, Dimitrios, et al.
Published: (2025)
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
by: Martinez, Matias
Published: (2024)
by: Martinez, Matias
Published: (2024)
LLM Critics Help Catch LLM Bugs
by: McAleese, Nat, et al.
Published: (2024)
by: McAleese, Nat, et al.
Published: (2024)
Real-Time Performance Benchmarking of TinyML Models in Embedded Systems (PICO: Performance of Inference, CPU, and Operations)
by: Dey, Abhishek, et al.
Published: (2025)
by: Dey, Abhishek, et al.
Published: (2025)
Mutation-Guided LLM-based Test Generation at Meta
by: Foster, Christopher, et al.
Published: (2025)
by: Foster, Christopher, et al.
Published: (2025)
Automating Formal Verification with Reinforcement Learning and Recursive Inference
by: Tan, Max
Published: (2026)
by: Tan, Max
Published: (2026)
Impact of ML Optimization Tactics on Greener Pre-Trained ML Models
by: Álvarez, Alexandra González, et al.
Published: (2024)
by: Álvarez, Alexandra González, et al.
Published: (2024)
Influence-Guided Concolic Testing of Transformer Robustness
by: Hong, Chih-Duo, et al.
Published: (2025)
by: Hong, Chih-Duo, et al.
Published: (2025)
RepDL: Bit-level Reproducible Deep Learning Training and Inference
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics
by: Patel, Disha
Published: (2026)
by: Patel, Disha
Published: (2026)
New Solutions on LLM Acceleration, Optimization, and Application
by: Huang, Yingbing, et al.
Published: (2024)
by: Huang, Yingbing, et al.
Published: (2024)
Grounded AI for Code Review: Resource-Efficient Large-Model Serving in Enterprise Pipelines
by: Mandal, Sayan, et al.
Published: (2025)
by: Mandal, Sayan, et al.
Published: (2025)
Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring
by: Ayotunde, John, et al.
Published: (2026)
by: Ayotunde, John, et al.
Published: (2026)
Rule-Guided Reinforcement Learning Policy Evaluation and Improvement
by: Tappler, Martin, et al.
Published: (2025)
by: Tappler, Martin, et al.
Published: (2025)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
LLM-Based Design Pattern Detection
by: Schindler, Christian, et al.
Published: (2025)
by: Schindler, Christian, et al.
Published: (2025)
Verifier-Guided Code Translation via Meta-Step Decoding
by: Zhou, Tianyang, et al.
Published: (2026)
by: Zhou, Tianyang, et al.
Published: (2026)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
by: Kim, Naryeong, et al.
Published: (2026)
by: Kim, Naryeong, et al.
Published: (2026)
Optimising for Energy Efficiency and Performance in Machine Learning
by: Ferreira, Emile Dos Santos, et al.
Published: (2026)
by: Ferreira, Emile Dos Santos, et al.
Published: (2026)
Similar Items
-
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
by: Dhulipala, Hridya, et al.
Published: (2025) -
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
by: Alizadeh, Negar, et al.
Published: (2024) -
Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters
by: Zine, Nada, et al.
Published: (2026) -
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
by: Prenner, Julian Aron, et al.
Published: (2025) -
Is Hyper-Parameter Optimization Different for Software Analytics?
by: Yedida, Rahul, et al.
Published: (2024)