Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Weihao, Zhang, Ji, Zhao, Han, Liu, Chao, Sha, Jian, He, Bingsheng, Guo, Minyi, Chen, Quan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)
by: Cai, Yifeng, et al.
Published: (2025)
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
by: Zhu, Hanqi, et al.
Published: (2025)
by: Zhu, Hanqi, et al.
Published: (2025)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
by: Shen, Weihang, et al.
Published: (2025)
by: Shen, Weihang, et al.
Published: (2025)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU
by: Sun, He, et al.
Published: (2025)
by: Sun, He, et al.
Published: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
On cleanness of AW*-algebras
by: Cui, Lu, et al.
Published: (2025)
by: Cui, Lu, et al.
Published: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024)
by: Wang, Tianxia, et al.
Published: (2024)
Live‐streaming selling with consumer returns: How does the sample strategy impact pricing decisions and streamer choice of the manufacturer?
by: Ping He, et al.
Published: (2026)
by: Ping He, et al.
Published: (2026)
Ensuring eco‐resilience in Physical Internet‐enabled production routing problem under ripple effect: a scenario‐based robust possibilistic flexible programming approach
by: Peng‐yun Zhao, et al.
Published: (2025)
by: Peng‐yun Zhao, et al.
Published: (2025)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
by: Zhang, Zongpu, et al.
Published: (2025)
by: Zhang, Zongpu, et al.
Published: (2025)
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
by: Han, Mingcong, et al.
Published: (2024)
by: Han, Mingcong, et al.
Published: (2024)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
RUISA Operational Ecosystem Architecture
by: AL Mohtar, Mouayad
Published: (2026)
by: AL Mohtar, Mouayad
Published: (2026)
Peformance Isolation for Inference Processes in Edge GPU Systems
by: Martín, Juan José, et al.
Published: (2026)
by: Martín, Juan José, et al.
Published: (2026)
Husky's Hidden Hand: A Compass for Change Agency and Subject Matter Expertise in Large Scale Combat Operations
by: Howard, Jeremy
Published: (2026)
by: Howard, Jeremy
Published: (2026)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
AgentRM: An OS-Inspired Resource Manager for LLM Agent Systems
by: She, Jianshu
Published: (2026)
by: She, Jianshu
Published: (2026)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026)
by: Yang, Zhenyuan, et al.
Published: (2026)
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026)
by: Zheng, Yusheng
Published: (2026)
Transition paths for condition‐based maintenance‐driven smart services
by: Henk Akkermans, et al.
Published: (2024)
by: Henk Akkermans, et al.
Published: (2024)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
What doesn't kill you makes you stronger? Evidence from vampire attacks on decentralized exchange and non‐fungible token marketplace
by: Xi Zhao, et al.
Published: (2024)
by: Xi Zhao, et al.
Published: (2024)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
by: Yang, Yiwei, et al.
Published: (2026)
by: Yang, Yiwei, et al.
Published: (2026)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
by: Dong, Yunpeng, et al.
Published: (2026)
by: Dong, Yunpeng, et al.
Published: (2026)
ByteFS: System Support for (CXL-based) Memory-Semantic Solid-State Drives
by: Li, Shaobo, et al.
Published: (2025)
by: Li, Shaobo, et al.
Published: (2025)
Similar Items
-
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025) -
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025) -
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026) -
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025) -
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)