EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jamil, Hasibul, Nine, MD S Q Zulkar, Kosar, Tevfik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Energy-Efficient and High-Performance Data Transfers with DRL Agents
von: Jamil, Hasibul, et al.
Veröffentlicht: (2025)
von: Jamil, Hasibul, et al.
Veröffentlicht: (2025)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
von: Jamil, Hasibul, et al.
Veröffentlicht: (2024)
von: Jamil, Hasibul, et al.
Veröffentlicht: (2024)
How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
Modeling Anomaly Detection in Cloud Services: Analysis of the Properties that Impact Latency and Resource Consumption
von: Grabher, Gabriel Job Antunes, et al.
Veröffentlicht: (2025)
von: Grabher, Gabriel Job Antunes, et al.
Veröffentlicht: (2025)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
von: Moran, Marina, et al.
Veröffentlicht: (2024)
von: Moran, Marina, et al.
Veröffentlicht: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
The Impact of Process Competition on Energy Consumption: Analysis and Modeling
von: Campos, Eduardo Gomes, et al.
Veröffentlicht: (2026)
von: Campos, Eduardo Gomes, et al.
Veröffentlicht: (2026)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Chasing the Speed of Light: Low-Latency Planetary-Scale Adaptive Byzantine Consensus
von: Berger, Christian, et al.
Veröffentlicht: (2023)
von: Berger, Christian, et al.
Veröffentlicht: (2023)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
Mobile Cloud Computing in Healthcare Using Dynamic Cloudlets for Energy-Aware Consumption
von: Muniswamaiah, Manoj, et al.
Veröffentlicht: (2019)
von: Muniswamaiah, Manoj, et al.
Veröffentlicht: (2019)
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
von: Gupta, Sushant Kumar, et al.
Veröffentlicht: (2025)
von: Gupta, Sushant Kumar, et al.
Veröffentlicht: (2025)
Falafels: A tool for Estimating Federated Learning Energy Consumption via Discrete Simulation
von: de Barochez, Andrew Mary Huet, et al.
Veröffentlicht: (2025)
von: de Barochez, Andrew Mary Huet, et al.
Veröffentlicht: (2025)
Strategies to Measure Energy Consumption Using RAPL During Workflow Execution on Commodity Clusters
von: Thamm, Philipp, et al.
Veröffentlicht: (2025)
von: Thamm, Philipp, et al.
Veröffentlicht: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
von: Pang, Zhiwen, et al.
Veröffentlicht: (2026)
von: Pang, Zhiwen, et al.
Veröffentlicht: (2026)
Enhancing Traffic Safety with AI and 6G: Latency Requirements and Real-Time Threat Detection
von: Horvath, Kurt, et al.
Veröffentlicht: (2025)
von: Horvath, Kurt, et al.
Veröffentlicht: (2025)
Understanding Power and Energy Utilization in Large Scale Production Physics Simulation Codes
von: Bertsch, Adam, et al.
Veröffentlicht: (2022)
von: Bertsch, Adam, et al.
Veröffentlicht: (2022)
Proof of Team Sprint: A Collaborative Consensus Algorithm for Reducing Energy Consumption in Blockchain Systems
von: Yonezawa, Naoki
Veröffentlicht: (2024)
von: Yonezawa, Naoki
Veröffentlicht: (2024)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
Minimizing Energy in Reliability and Deadline-Ensured Workflow Scheduling in Cloud
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2025)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2025)
Asynchronous Latency and Fast Atomic Snapshot
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
HE2C: A Holistic Approach for Allocating Latency-Sensitive AI Tasks across Edge-Cloud
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
A Periodic Space of Distributed Computing: Vision & Framework
von: Salehi, Mohsen Amini, et al.
Veröffentlicht: (2026)
von: Salehi, Mohsen Amini, et al.
Veröffentlicht: (2026)
Large-Scale Graph Building in Dynamic Environments: Low Latency and High Quality
von: de Almeida, Filipe Miguel Gonçalves, et al.
Veröffentlicht: (2025)
von: de Almeida, Filipe Miguel Gonçalves, et al.
Veröffentlicht: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
DAG it off: Latency Prefers No Common Coins
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
Methodology for GPU Frequency Switching Latency Measurement
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
Redox: Improving I/O Efficiency of Model Training Through File Redirection
von: Li, Yuhao, et al.
Veröffentlicht: (2025)
von: Li, Yuhao, et al.
Veröffentlicht: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models
von: Rajgopal, Ajay Navilarekal, et al.
Veröffentlicht: (2026)
von: Rajgopal, Ajay Navilarekal, et al.
Veröffentlicht: (2026)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Energy-Efficient and High-Performance Data Transfers with DRL Agents
von: Jamil, Hasibul, et al.
Veröffentlicht: (2025) -
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
von: Niam, Arefin, et al.
Veröffentlicht: (2026) -
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025) -
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
von: Jamil, Hasibul, et al.
Veröffentlicht: (2024) -
How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)