AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Ahan, Wang, Zhihao, Dani, Neel, Tanaka, Masahiro, Ruwase, Olatunji, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
MegaFold: System-Level Optimizations for Accelerating Protein Structure Prediction Models
von: La, Hoa, et al.
Veröffentlicht: (2025)
von: La, Hoa, et al.
Veröffentlicht: (2025)
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
Kino-PAX: Highly Parallel Kinodynamic Sampling-based Planner
von: Perrault, Nicolas, et al.
Veröffentlicht: (2024)
von: Perrault, Nicolas, et al.
Veröffentlicht: (2024)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2026)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2026)
SP-IMPact: A Framework for Static Partitioning Interference Mitigation and Performance Analysis
von: Costa, Diogo, et al.
Veröffentlicht: (2025)
von: Costa, Diogo, et al.
Veröffentlicht: (2025)
Comparing Parallel Functional Array Languages: Programming and Performance
von: van Balen, David, et al.
Veröffentlicht: (2025)
von: van Balen, David, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025) -
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025) -
MegaFold: System-Level Optimizations for Accelerating Protein Structure Prediction Models
von: La, Hoa, et al.
Veröffentlicht: (2025) -
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024) -
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)