CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913816840765440 |
|---|---|
| author | Paramanayakam, Varatheepan Karatzas, Andreas Anagnostopoulos, Iraklis Stamoulis, Dimitrios |
| author_facet | Paramanayakam, Varatheepan Karatzas, Andreas Anagnostopoulos, Iraklis Stamoulis, Dimitrios |
| contents | Large Language Models (LLMs) enable real-time function calling in edge AI systems but introduce significant computational overhead, leading to high power consumption and carbon emissions. Existing methods optimize for performance while neglecting sustainability, making them inefficient for energy-constrained environments. We introduce CarbonCall, a sustainability-aware function-calling framework that integrates dynamic tool selection, carbon-aware execution, and quantized LLM adaptation. CarbonCall adjusts power thresholds based on real-time carbon intensity forecasts and switches between model variants to sustain high tokens-per-second throughput under power constraints. Experiments on an NVIDIA Jetson AGX Orin show that CarbonCall reduces carbon emissions by up to 52%, power consumption by 30%, and execution time by 30%, while maintaining high efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_20348 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices Paramanayakam, Varatheepan Karatzas, Andreas Anagnostopoulos, Iraklis Stamoulis, Dimitrios Performance Artificial Intelligence Systems and Control Large Language Models (LLMs) enable real-time function calling in edge AI systems but introduce significant computational overhead, leading to high power consumption and carbon emissions. Existing methods optimize for performance while neglecting sustainability, making them inefficient for energy-constrained environments. We introduce CarbonCall, a sustainability-aware function-calling framework that integrates dynamic tool selection, carbon-aware execution, and quantized LLM adaptation. CarbonCall adjusts power thresholds based on real-time carbon intensity forecasts and switches between model variants to sustain high tokens-per-second throughput under power constraints. Experiments on an NVIDIA Jetson AGX Orin show that CarbonCall reduces carbon emissions by up to 52%, power consumption by 30%, and execution time by 30%, while maintaining high efficiency. |
| title | CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices |
| topic | Performance Artificial Intelligence Systems and Control |
| url | https://arxiv.org/abs/2504.20348 |