Hierarchical Prediction-based Management for LMaaS Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Zhihan, Huang, Yujie, Yu, Guangba, Huang, Junjie, Gu, Jiazhen, Lyu, Michael R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
di: Wang, Zirui, et al.
Pubblicazione: (2026)
di: Wang, Zirui, et al.
Pubblicazione: (2026)
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
di: Yu, Guangba, et al.
Pubblicazione: (2026)
di: Yu, Guangba, et al.
Pubblicazione: (2026)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
di: Yu, Guangba, et al.
Pubblicazione: (2026)
di: Yu, Guangba, et al.
Pubblicazione: (2026)
sVIRGO: A Scalable Virtual Tree Hierarchical Framework for Distributed Systems
di: Huang, Lican
Pubblicazione: (2026)
di: Huang, Lican
Pubblicazione: (2026)
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
A Survey on Failure Analysis and Fault Injection in AI Systems
di: Yu, Guangba, et al.
Pubblicazione: (2024)
di: Yu, Guangba, et al.
Pubblicazione: (2024)
Designing Co-operation in Systems of Hierarchical, Multi-objective Schedulers for Stream Processing
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
Efficient Hierarchical Storage Management Framework Empowered by Reinforcement Learning
di: Zhang, Tianru, et al.
Pubblicazione: (2022)
di: Zhang, Tianru, et al.
Pubblicazione: (2022)
Coordinated Power Management on Heterogeneous Systems
di: Zheng, Zhong, et al.
Pubblicazione: (2025)
di: Zheng, Zhong, et al.
Pubblicazione: (2025)
Adaptive Cache Management for Complex Storage Systems Using CNN-LSTM-Based Spatiotemporal Prediction
di: Wang, Xiaoye, et al.
Pubblicazione: (2024)
di: Wang, Xiaoye, et al.
Pubblicazione: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
di: Zhu, Jianyong, et al.
Pubblicazione: (2026)
di: Zhu, Jianyong, et al.
Pubblicazione: (2026)
TSUE: A Two-Stage Data Update Method for an Erasure Coded Cluster File System
di: Wei, Zheng, et al.
Pubblicazione: (2025)
di: Wei, Zheng, et al.
Pubblicazione: (2025)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
di: Zheng, Zhong, et al.
Pubblicazione: (2026)
di: Zheng, Zhong, et al.
Pubblicazione: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
Graph for Science: From API based Programming to Graph Engine based Programming for HPC
di: Zhang, Yu, et al.
Pubblicazione: (2023)
di: Zhang, Yu, et al.
Pubblicazione: (2023)
Squeezing Edge Performance: A Sensitivity-Aware Container Management for Heterogeneous Tasks
di: Zhang, Yongmin, et al.
Pubblicazione: (2025)
di: Zhang, Yongmin, et al.
Pubblicazione: (2025)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
di: Zhong, Xinrui, et al.
Pubblicazione: (2025)
di: Zhong, Xinrui, et al.
Pubblicazione: (2025)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
CodeAD: Synthesize Code of Rules for Log-based Anomaly Detection with LLMs
di: Huang, Junjie, et al.
Pubblicazione: (2025)
di: Huang, Junjie, et al.
Pubblicazione: (2025)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
di: Wang, Yuxiang, et al.
Pubblicazione: (2025)
di: Wang, Yuxiang, et al.
Pubblicazione: (2025)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
Pruning Blockchain Protocols for Efficient Access Control in IoT Systems
di: Huang, Yongtao, et al.
Pubblicazione: (2024)
di: Huang, Yongtao, et al.
Pubblicazione: (2024)
Development of a Cloud-Based Payroll Management System
di: Aina, Adeyemi, et al.
Pubblicazione: (2025)
di: Aina, Adeyemi, et al.
Pubblicazione: (2025)
Data Management System Analysis for Distributed Computing Workloads
di: Hsu, Kuan-Chieh, et al.
Pubblicazione: (2025)
di: Hsu, Kuan-Chieh, et al.
Pubblicazione: (2025)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
di: Rashid, Sulaiman Muhammad, et al.
Pubblicazione: (2025)
di: Rashid, Sulaiman Muhammad, et al.
Pubblicazione: (2025)
PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
di: Olama, Alireza, et al.
Pubblicazione: (2025)
di: Olama, Alireza, et al.
Pubblicazione: (2025)
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
di: Fang, Zhengxin, et al.
Pubblicazione: (2025)
di: Fang, Zhengxin, et al.
Pubblicazione: (2025)
RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
di: Gener, Serhan, et al.
Pubblicazione: (2025)
di: Gener, Serhan, et al.
Pubblicazione: (2025)
Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art
di: Bader, Jonathan, et al.
Pubblicazione: (2025)
di: Bader, Jonathan, et al.
Pubblicazione: (2025)
Tracing Cross-chain Transactions between EVM-based Blockchains: An Analysis of Ethereum-Polygon Bridges
di: Yan, Tao, et al.
Pubblicazione: (2025)
di: Yan, Tao, et al.
Pubblicazione: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
Exploring the Frontiers of Energy Efficiency using Power Management at System Scale
di: Karimi, Ahmad Maroof, et al.
Pubblicazione: (2024)
di: Karimi, Ahmad Maroof, et al.
Pubblicazione: (2024)
Managing Forensic Recovery in the Cloud
di: Weir, George R. S., et al.
Pubblicazione: (2024)
di: Weir, George R. S., et al.
Pubblicazione: (2024)
Tolerating Disasters with Hierarchical Consensus
di: Yahyaoui, Wassim, et al.
Pubblicazione: (2025)
di: Yahyaoui, Wassim, et al.
Pubblicazione: (2025)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
di: Yuan, Hao, et al.
Pubblicazione: (2023)
di: Yuan, Hao, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
di: Cheng, Ke, et al.
Pubblicazione: (2024) -
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025) -
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
di: Wang, Zirui, et al.
Pubblicazione: (2026) -
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
di: Yu, Guangba, et al.
Pubblicazione: (2026) -
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
di: Yu, Guangba, et al.
Pubblicazione: (2026)