SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jaiswal, Shashwat, Jain, Kunal, Simmhan, Yogesh, Parayil, Anjaly, Mallick, Ankur, Wang, Rujia, Amant, Renee St., Bansal, Chetan, Rühle, Victor, Kulkarni, Anoop, Kofsky, Steve, Rajmohan, Saravan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ensuring Fair LLM Serving Amid Diverse Applications
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
von: Hussain, Fiza, et al.
Veröffentlicht: (2025)
von: Hussain, Fiza, et al.
Veröffentlicht: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
von: Srinivas, Pooja, et al.
Veröffentlicht: (2024)
von: Srinivas, Pooja, et al.
Veröffentlicht: (2024)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
von: Bastos, Anson, et al.
Veröffentlicht: (2026)
von: Bastos, Anson, et al.
Veröffentlicht: (2026)
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
X-lifecycle Learning for Cloud Incident Management using LLMs
von: Goel, Drishti, et al.
Veröffentlicht: (2024)
von: Goel, Drishti, et al.
Veröffentlicht: (2024)
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
von: Xia, Menglin, et al.
Veröffentlicht: (2026)
von: Xia, Menglin, et al.
Veröffentlicht: (2026)
AerialDB: A Federated Peer-to-Peer Spatio-temporal Edge Datastore for Drone Fleets
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
AutoAdapt: An Automated Domain Adaptation Framework for LLMs
von: Sinha, Sidharth, et al.
Veröffentlicht: (2026)
von: Sinha, Sidharth, et al.
Veröffentlicht: (2026)
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4
von: Zhang, Xuchao, et al.
Veröffentlicht: (2024)
von: Zhang, Xuchao, et al.
Veröffentlicht: (2024)
Towards AI Agents for Course Instruction in Higher Education: Early Experiences from the Field
von: Simmhan, Yogesh, et al.
Veröffentlicht: (2025)
von: Simmhan, Yogesh, et al.
Veröffentlicht: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Exploring How LLMs Capture and Represent Domain-Specific Knowledge
von: Garcia, Mirian Hipolito, et al.
Veröffentlicht: (2025)
von: Garcia, Mirian Hipolito, et al.
Veröffentlicht: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
von: Xu, Jin, et al.
Veröffentlicht: (2026)
von: Xu, Jin, et al.
Veröffentlicht: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
von: Couturier, Camille, et al.
Veröffentlicht: (2025)
von: Couturier, Camille, et al.
Veröffentlicht: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
Gaussian beams and inverse problems for connections at high fixed frequency
von: St-Amant, Simon
Veröffentlicht: (2024)
von: St-Amant, Simon
Veröffentlicht: (2024)
The Gulf States Marine Fisheries Commission and the Menhaden Fishery
von: St. Amant, L.S.
Veröffentlicht: (1974)
von: St. Amant, L.S.
Veröffentlicht: (1974)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
von: Xia, Menglin, et al.
Veröffentlicht: (2023)
von: Xia, Menglin, et al.
Veröffentlicht: (2023)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
Dependency Aware Incident Linking in Large Cloud Systems
von: Ghosh, Supriyo, et al.
Veröffentlicht: (2024)
von: Ghosh, Supriyo, et al.
Veröffentlicht: (2024)
Building AI Agents for Autonomous Clouds: Challenges and Design Principles
von: Shetty, Manish, et al.
Veröffentlicht: (2024)
von: Shetty, Manish, et al.
Veröffentlicht: (2024)
LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
von: Han, Dongge, et al.
Veröffentlicht: (2025)
von: Han, Dongge, et al.
Veröffentlicht: (2025)
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
von: Han, Dongge, et al.
Veröffentlicht: (2025)
von: Han, Dongge, et al.
Veröffentlicht: (2025)
Exploring LLM-based Agents for Root Cause Analysis
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
SynthAgent: Adapting Web Agents with Synthetic Supervision
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
Budget-Aware Agentic Routing via Boundary-Guided Training
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ensuring Fair LLM Serving Amid Diverse Applications
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024) -
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025) -
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024) -
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
von: Hussain, Fiza, et al.
Veröffentlicht: (2025) -
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)