GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ilager, Shashikant, Briem, Lukas Florian, Brandic, Ivona |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
par: May, Daniel, et autres
Publié: (2024)
par: May, Daniel, et autres
Publié: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
par: Kolluru, Saicharan
Publié: (2025)
par: Kolluru, Saicharan
Publié: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
par: Topcu, Burak, et autres
Publié: (2026)
par: Topcu, Burak, et autres
Publié: (2026)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
par: Zilic, Josip, et autres
Publié: (2024)
par: Zilic, Josip, et autres
Publié: (2024)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
par: Nichols, Daniel, et autres
Publié: (2025)
par: Nichols, Daniel, et autres
Publié: (2025)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
par: Ataie, Ehsan, et autres
Publié: (2026)
par: Ataie, Ehsan, et autres
Publié: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
par: Jo, Myeong Jun
Publié: (2026)
par: Jo, Myeong Jun
Publié: (2026)
ParaQAOA: Efficient Parallel Divide-and-Conquer QAOA for Large-Scale Max-Cut Problems Beyond 10,000 Vertices
par: Huang, Po-Hsuan, et autres
Publié: (2026)
par: Huang, Po-Hsuan, et autres
Publié: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
par: Wang, Zirui, et autres
Publié: (2026)
par: Wang, Zirui, et autres
Publié: (2026)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
par: Vogel, Adriano, et autres
Publié: (2024)
par: Vogel, Adriano, et autres
Publié: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
par: Tariq, Syed Salauddin Mohammad, et autres
Publié: (2024)
par: Tariq, Syed Salauddin Mohammad, et autres
Publié: (2024)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
par: Henning, Sören, et autres
Publié: (2025)
par: Henning, Sören, et autres
Publié: (2025)
Where Should I Deploy My Contracts? A Practical Experience Report
par: Lazăr, Cătălina, et autres
Publié: (2025)
par: Lazăr, Cătălina, et autres
Publié: (2025)
MPI Implementation Profiling for Better Application Performance
par: Shipley, Riley, et autres
Publié: (2024)
par: Shipley, Riley, et autres
Publié: (2024)
Should I Run My Cloud Benchmark on Black Friday?
par: Henning, Sören, et autres
Publié: (2025)
par: Henning, Sören, et autres
Publié: (2025)
Intent-driven scheduling of backup jobs
par: Dutta, Souvik, et autres
Publié: (2024)
par: Dutta, Souvik, et autres
Publié: (2024)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
par: Chen, Mu-Chi, et autres
Publié: (2025)
par: Chen, Mu-Chi, et autres
Publié: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
par: Lei, Jianlong, et autres
Publié: (2026)
par: Lei, Jianlong, et autres
Publié: (2026)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
par: Kamath, Aditya K, et autres
Publié: (2026)
par: Kamath, Aditya K, et autres
Publié: (2026)
Emergence-as-Code for Self-Governing Reliable Systems
par: Krasnovsky, Anatoly A.
Publié: (2026)
par: Krasnovsky, Anatoly A.
Publié: (2026)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
par: Kharma, Sami, et autres
Publié: (2026)
par: Kharma, Sami, et autres
Publié: (2026)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
par: Miletto, Marcelo Cogo, et autres
Publié: (2024)
par: Miletto, Marcelo Cogo, et autres
Publié: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
par: Tuteja, Keshvi, et autres
Publié: (2025)
par: Tuteja, Keshvi, et autres
Publié: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
par: Li, Junjie
Publié: (2024)
par: Li, Junjie
Publié: (2024)
Serverless Cold Starts and Where to Find Them
par: Joosen, Artjom, et autres
Publié: (2024)
par: Joosen, Artjom, et autres
Publié: (2024)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
par: Ganjihal, Sanjeev Rao
Publié: (2026)
par: Ganjihal, Sanjeev Rao
Publié: (2026)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
par: He, Jiaao, et autres
Publié: (2024)
par: He, Jiaao, et autres
Publié: (2024)
Designing Scalable Rate Limiting Systems: Algorithms, Architecture, and Distributed Solutions
par: Guan, Bo
Publié: (2026)
par: Guan, Bo
Publié: (2026)
Efficient Construction of Large Search Spaces for Auto-Tuning
par: Willemsen, Floris-Jan, et autres
Publié: (2025)
par: Willemsen, Floris-Jan, et autres
Publié: (2025)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
par: Ruiz, Christian Guzman, et autres
Publié: (2024)
par: Ruiz, Christian Guzman, et autres
Publié: (2024)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
par: Yuan, Renzhong, et autres
Publié: (2026)
par: Yuan, Renzhong, et autres
Publié: (2026)
A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
par: Ilager, Shashikant, et autres
Publié: (2024)
par: Ilager, Shashikant, et autres
Publié: (2024)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
par: Wen, Zhongzhen, et autres
Publié: (2025)
par: Wen, Zhongzhen, et autres
Publié: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
par: Singh, Samaresh Kumar, et autres
Publié: (2026)
par: Singh, Samaresh Kumar, et autres
Publié: (2026)
Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective
par: Huang, Mingyu, et autres
Publié: (2024)
par: Huang, Mingyu, et autres
Publié: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
par: Krasnovsky, Anatoly A.
Publié: (2025)
par: Krasnovsky, Anatoly A.
Publié: (2025)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
par: Wen, Zhongzhen, et autres
Publié: (2026)
par: Wen, Zhongzhen, et autres
Publié: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
par: Malhotra, Shubham, et autres
Publié: (2025)
par: Malhotra, Shubham, et autres
Publié: (2025)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
par: Teranishi, Keita, et autres
Publié: (2025)
par: Teranishi, Keita, et autres
Publié: (2025)
Documents similaires
-
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
par: May, Daniel, et autres
Publié: (2024) -
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
par: Kolluru, Saicharan
Publié: (2025) -
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
par: Topcu, Burak, et autres
Publié: (2026) -
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
par: Zilic, Josip, et autres
Publié: (2024) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
par: Nichols, Daniel, et autres
Publié: (2025)