SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qu, Lei, Ren, Lianhai, Cheng, Peng, Gao, Rui, Wang, Ruizhe, Chen, Tianyu, Liu, Xiao, Zhang, Xingjian, Gong, Yeyun, Xiong, Yifan, Ding, Yucheng, Jiang, Yuting, Lin, Zhenghao, Guo, Zhongxin, Yang, Ziyue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
On Software Ageing Indicators in OpenStack
von: Yazvinskyi, Yevhen, et al.
Veröffentlicht: (2024)
von: Yazvinskyi, Yevhen, et al.
Veröffentlicht: (2024)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
HopGNN: Boosting Distributed GNN Training Efficiency via Feature-Centric Model Migration
von: Chen, Weijian, et al.
Veröffentlicht: (2024)
von: Chen, Weijian, et al.
Veröffentlicht: (2024)
Empowering the Quantum Cloud User with QRIO
von: Chakraborty, Shmeelok, et al.
Veröffentlicht: (2024)
von: Chakraborty, Shmeelok, et al.
Veröffentlicht: (2024)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026)
von: Ma, Bin, et al.
Veröffentlicht: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
Enabling an OpenStack-based cloud on top of RISC-V hardware
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
von: Gong, Yidong, et al.
Veröffentlicht: (2026)
von: Gong, Yidong, et al.
Veröffentlicht: (2026)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
von: Navardi, Mozhgan, et al.
Veröffentlicht: (2025)
von: Navardi, Mozhgan, et al.
Veröffentlicht: (2025)
Efficient Hierarchical Storage Management Framework Empowered by Reinforcement Learning
von: Zhang, Tianru, et al.
Veröffentlicht: (2022)
von: Zhang, Tianru, et al.
Veröffentlicht: (2022)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)
von: Shubham, et al.
Veröffentlicht: (2024)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
HMTRace: Hardware-Assisted Memory-Tagging based Dynamic Data Race Detection
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
von: Takahashi, Keichi, et al.
Veröffentlicht: (2024)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2024)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
Comparing the Run-time Behavior of Modern PDES Engines on Alternative Hardware Architectures
von: Marotta, Romolo, et al.
Veröffentlicht: (2025)
von: Marotta, Romolo, et al.
Veröffentlicht: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
Empowering Distributed Training with Sparsity-driven Data Synchronization
von: Wang, Zhuang, et al.
Veröffentlicht: (2023)
von: Wang, Zhuang, et al.
Veröffentlicht: (2023)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024) -
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
von: Xiong, Yifan, et al.
Veröffentlicht: (2024) -
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026) -
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024) -
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)