Efficient Multi-Model Orchestration for Self-Hosted Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vangala, Bhanu Prakash, Malik, Tanu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows
von: Ifath, Md. Monzurul Amin, et al.
Veröffentlicht: (2026)
von: Ifath, Md. Monzurul Amin, et al.
Veröffentlicht: (2026)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
Scaling Performance of Large Language Model Pretraining
von: Interrante-Grant, Alexander, et al.
Veröffentlicht: (2025)
von: Interrante-Grant, Alexander, et al.
Veröffentlicht: (2025)
HPC-Coder: Modeling Parallel Programs using Large Language Models
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
Can Large Language Models Write Parallel Code?
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)
von: He, Zifan, et al.
Veröffentlicht: (2026)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
Research on Model Parallelism and Data Parallelism Optimization Methods in Large Language Model-Based Recommendation Systems
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
Adaptive Fault Tolerance Mechanisms of Large Language Models in Cloud Computing Environments
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
von: Han, Shujie, et al.
Veröffentlicht: (2026)
von: Han, Shujie, et al.
Veröffentlicht: (2026)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
A Survey on Large Language Model Acceleration based on KV Cache Management
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
von: Huang, Yafan, et al.
Veröffentlicht: (2026)
von: Huang, Yafan, et al.
Veröffentlicht: (2026)
Connecting Large Language Models with Blockchain: Advancing the Evolution of Smart Contracts from Automation to Intelligence
von: Xian, Youquan, et al.
Veröffentlicht: (2024)
von: Xian, Youquan, et al.
Veröffentlicht: (2024)
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
von: Lu, Ning, et al.
Veröffentlicht: (2024)
von: Lu, Ning, et al.
Veröffentlicht: (2024)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
von: Fang, Fei, et al.
Veröffentlicht: (2025)
von: Fang, Fei, et al.
Veröffentlicht: (2025)
Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
Efficient Federated Fine-Tuning of Large Language Models with Layer Dropout
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
Can Large Language Models Predict Parallel Code Performance?
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models
von: Zhang, Zishuai, et al.
Veröffentlicht: (2025)
von: Zhang, Zishuai, et al.
Veröffentlicht: (2025)
A Hashgraph-Inspired Consensus Mechanism for Reliable Multi-Model Reasoning
von: Ogunsina, Kolawole E., et al.
Veröffentlicht: (2025)
von: Ogunsina, Kolawole E., et al.
Veröffentlicht: (2025)
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025) -
Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows
von: Ifath, Md. Monzurul Amin, et al.
Veröffentlicht: (2026) -
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025) -
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024) -
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
von: Ji, Cheng, et al.
Veröffentlicht: (2025)