Beyond Efficiency: Scaling AI Sustainably
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Carole-Jean, Acun, Bilge, Raghavendra, Ramya, Hazelwood, Kim |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
by: Golden, Alicia, et al.
Published: (2023)
by: Golden, Alicia, et al.
Published: (2023)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023)
by: Hsia, Samuel, et al.
Published: (2023)
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
by: Kim, Gyudong, et al.
Published: (2024)
by: Kim, Gyudong, et al.
Published: (2024)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
by: Oviedo, Felipe, et al.
Published: (2025)
by: Oviedo, Felipe, et al.
Published: (2025)
Beyond Model Scale Limits: End-Edge-Cloud Federated Learning with Self-Rectified Knowledge Agglomeration
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Intelligent Task Offloading in VANETs: A Hybrid AI-Driven Approach for Low-Latency and Energy Efficiency
by: Qayyum, Tariq, et al.
Published: (2025)
by: Qayyum, Tariq, et al.
Published: (2025)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
ReInc: Scaling Training of Dynamic Graph Neural Networks
by: Guan, Mingyu, et al.
Published: (2025)
by: Guan, Mingyu, et al.
Published: (2025)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
by: Tschand, Arya, et al.
Published: (2024)
by: Tschand, Arya, et al.
Published: (2024)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
by: Zhu, Ruidong, et al.
Published: (2025)
by: Zhu, Ruidong, et al.
Published: (2025)
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
by: Wang, Weixun, et al.
Published: (2025)
by: Wang, Weixun, et al.
Published: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
by: Tian, Jian, et al.
Published: (2025)
by: Tian, Jian, et al.
Published: (2025)
EMO: Edge Model Overlays to Scale Model Size in Federated Learning
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Towards Sustainable Large Language Model Serving
by: Nguyen, Sophia, et al.
Published: (2024)
by: Nguyen, Sophia, et al.
Published: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
by: Chai, Huichao, et al.
Published: (2026)
by: Chai, Huichao, et al.
Published: (2026)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
by: Pan, Yi, et al.
Published: (2025)
by: Pan, Yi, et al.
Published: (2025)
Peer-to-Peer Deep Learning for Beyond-5G IoT
by: Pranav, Srinivasa, et al.
Published: (2023)
by: Pranav, Srinivasa, et al.
Published: (2023)
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
by: Dong, Jianbo, et al.
Published: (2024)
by: Dong, Jianbo, et al.
Published: (2024)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
by: Malekzadeh, Mohammad, et al.
Published: (2023)
by: Malekzadeh, Mohammad, et al.
Published: (2023)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
by: Jiang, Ziheng, et al.
Published: (2024)
by: Jiang, Ziheng, et al.
Published: (2024)
Beyond the Federation: Topology-aware Federated Learning for Generalization to Unseen Clients
by: Ma, Mengmeng, et al.
Published: (2024)
by: Ma, Mengmeng, et al.
Published: (2024)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
by: Shen, Ao, et al.
Published: (2024)
by: Shen, Ao, et al.
Published: (2024)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
by: Fang, Jingzhi, et al.
Published: (2025)
by: Fang, Jingzhi, et al.
Published: (2025)
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
Synergy: Towards On-Body AI via Tiny AI Accelerator Collaboration on Wearables
by: Gong, Taesik, et al.
Published: (2023)
by: Gong, Taesik, et al.
Published: (2023)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
by: Liu, Dennis, et al.
Published: (2025)
by: Liu, Dennis, et al.
Published: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry
by: Bidollahkhani, Michael, et al.
Published: (2026)
by: Bidollahkhani, Michael, et al.
Published: (2026)
Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning
by: Budiarjo, Thomas, et al.
Published: (2025)
by: Budiarjo, Thomas, et al.
Published: (2025)
Provenance Tracking in Large-Scale Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Similar Items
-
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
by: Golden, Alicia, et al.
Published: (2023) -
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023) -
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024) -
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
by: Kim, Gyudong, et al.
Published: (2024) -
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)