CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Paramanayakam, Varatheepan, Karatzas, Andreas, Anagnostopoulos, Iraklis, Stamoulis, Dimitrios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
by: Paramanayakam, Varatheepan, et al.
Published: (2024)
by: Paramanayakam, Varatheepan, et al.
Published: (2024)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
A Vertical Approach to Designing and Managing Sustainable Heterogeneous Edge Data Centers
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded Devices
by: Karatzas, Andreas, et al.
Published: (2024)
by: Karatzas, Andreas, et al.
Published: (2024)
Revolutionizing System Reliability: The Role of AI in Predictive Maintenance Strategies
by: Bidollahkhani, Michael, et al.
Published: (2024)
by: Bidollahkhani, Michael, et al.
Published: (2024)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
The Unseen AI Disruptions for Power Grids: LLM-Induced Transients
by: Li, Yuzhuo, et al.
Published: (2024)
by: Li, Yuzhuo, et al.
Published: (2024)
Performance Prediction for Large Systems via Text-to-Text Regression
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
by: Badmus, Emmanuel O., et al.
Published: (2025)
by: Badmus, Emmanuel O., et al.
Published: (2025)
QuickGrasp: Responsive Video-Language Querying Service via Accelerated Tokenization and Edge-Augmented Inference
by: Zhang, Miao, et al.
Published: (2026)
by: Zhang, Miao, et al.
Published: (2026)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)
by: Taufique, Zain, et al.
Published: (2023)
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
by: Yu, Nengneng, et al.
Published: (2026)
by: Yu, Nengneng, et al.
Published: (2026)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
by: Hossain, Abrar, et al.
Published: (2025)
by: Hossain, Abrar, et al.
Published: (2025)
Multi-Agent Geospatial Copilots for Remote Sensing Workflows
by: Lee, Chaehong, et al.
Published: (2025)
by: Lee, Chaehong, et al.
Published: (2025)
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum
by: De Silva, Suvi, et al.
Published: (2026)
by: De Silva, Suvi, et al.
Published: (2026)
Every Call is Precious: Global Optimization of Black-Box Functions with Unknown Lipschitz Constants
by: Fourati, Fares, et al.
Published: (2025)
by: Fourati, Fares, et al.
Published: (2025)
Turning AI Data Centers into Grid-Interactive Assets: Results from a Field Demonstration in Phoenix, Arizona
by: Colangelo, Philip, et al.
Published: (2025)
by: Colangelo, Philip, et al.
Published: (2025)
Deep Learning for Low-Latency, Quantum-Ready RF Sensing
by: Gokhale, Pranav, et al.
Published: (2024)
by: Gokhale, Pranav, et al.
Published: (2024)
Unleashing Automated Congestion Control Customization in the Wild
by: Cohen, Amit, et al.
Published: (2025)
by: Cohen, Amit, et al.
Published: (2025)
Selecting Offline Reinforcement Learning Algorithms for Stochastic Network Control
by: Helson, Nicolas, et al.
Published: (2026)
by: Helson, Nicolas, et al.
Published: (2026)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
by: Hoefler, Torsten, et al.
Published: (2026)
by: Hoefler, Torsten, et al.
Published: (2026)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
by: Chai, Yuji, et al.
Published: (2025)
by: Chai, Yuji, et al.
Published: (2025)
Sustainable Edge Intelligence Through Energy-Aware Early Exiting
by: Bullo, Marcello, et al.
Published: (2023)
by: Bullo, Marcello, et al.
Published: (2023)
On the Sustainability of AI Inferences in the Edge
by: Sobhani, Ghazal, et al.
Published: (2025)
by: Sobhani, Ghazal, et al.
Published: (2025)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
AI-Driven Resource Allocation Framework for Microservices in Hybrid Cloud Platforms
by: Barua, Biman, et al.
Published: (2024)
by: Barua, Biman, et al.
Published: (2024)
Equilibrium in the Computing Continuum through Active Inference
by: Sedlak, Boris, et al.
Published: (2023)
by: Sedlak, Boris, et al.
Published: (2023)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
by: Prieto, Pablo, et al.
Published: (2025)
by: Prieto, Pablo, et al.
Published: (2025)
Modeling and Controlling Many-Core HPC Processors: an Alternative to PID and Moving Average Algorithms
by: Bambini, Giovanni, et al.
Published: (2024)
by: Bambini, Giovanni, et al.
Published: (2024)
Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration
by: Ou, Ninglin, et al.
Published: (2026)
by: Ou, Ninglin, et al.
Published: (2026)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Muninn: Your Trajectory Diffusion Model But Faster
by: Puthumanaillam, Gokul, et al.
Published: (2026)
by: Puthumanaillam, Gokul, et al.
Published: (2026)
High-performance computing enabled contingency analysis for modern power networks
by: Gracia-Calvo, Alexandre, et al.
Published: (2025)
by: Gracia-Calvo, Alexandre, et al.
Published: (2025)
Emergency Department Patient Flow Optimization with an Alternative Care Threshold Policy
by: Baniasadi, Sahba, et al.
Published: (2026)
by: Baniasadi, Sahba, et al.
Published: (2026)
Distributed Tracing for Cascading Changes of Objects in the Kubernetes Control Plane
by: Ehira, Tomoyuki, et al.
Published: (2024)
by: Ehira, Tomoyuki, et al.
Published: (2024)
Balanced allocation: considerations from large scale service environments
by: Diwan, Amer, et al.
Published: (2026)
by: Diwan, Amer, et al.
Published: (2026)
Exploiting Scheduling Flexibility via State-Based Scheduling When Guaranteeing Worst-Case Services
by: Xu, Yike, et al.
Published: (2026)
by: Xu, Yike, et al.
Published: (2026)
Similar Items
-
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
by: Paramanayakam, Varatheepan, et al.
Published: (2024) -
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025) -
A Vertical Approach to Designing and Managing Sustainable Heterogeneous Edge Data Centers
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025) -
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024) -
RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded Devices
by: Karatzas, Andreas, et al.
Published: (2024)