Gespeichert in:
| Hauptverfasser: | Hakim, Sheikh Azizul, Hasan, Saem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.11211 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Galvatron: Automatic Distributed Training for Large Transformer Models
von: Gumaan, Esmail
Veröffentlicht: (2025)
von: Gumaan, Esmail
Veröffentlicht: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
von: Kundu, Joyjit, et al.
Veröffentlicht: (2024)
von: Kundu, Joyjit, et al.
Veröffentlicht: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
Inference Acceleration for Large Language Models on CPUs
von: PS, Ditto, et al.
Veröffentlicht: (2024)
von: PS, Ditto, et al.
Veröffentlicht: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
Comparative Study of Large Language Model Architectures on Frontier
von: Yin, Junqi, et al.
Veröffentlicht: (2024)
von: Yin, Junqi, et al.
Veröffentlicht: (2024)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
Proactive and Reactive Autoscaling Techniques for Edge Computing
von: Gupta, Suhrid, et al.
Veröffentlicht: (2025)
von: Gupta, Suhrid, et al.
Veröffentlicht: (2025)
AI Surrogate Model for Distributed Computing Workloads
von: Park, David K., et al.
Veröffentlicht: (2024)
von: Park, David K., et al.
Veröffentlicht: (2024)
TF-DDRL: A Transformer-enhanced Distributed DRL Technique for Scheduling IoT Applications in Edge and Cloud Computing Environments
von: Wang, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyu, et al.
Veröffentlicht: (2024)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
A Study on the Performance of Distributed Training of Data-driven CFD Simulations
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026) -
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025) -
Characterizing Communication Patterns in Distributed Large Language Model Inference
von: Xu, Lang, et al.
Veröffentlicht: (2025) -
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)