GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
Fuente:
arXiv
Salvato in:
| Autori principali: | Jeon, Byungsoo, Wu, Mengdi, Cao, Shiyi, Kim, Sunghyun, Park, Sunghyun, Aggarwal, Neeraj, Unger, Colin, Arfeen, Daiyaan, Liao, Peiyuan, Miao, Xupeng, Alizadeh, Mohammad, Ganger, Gregory R., Chen, Tianqi, Jia, Zhihao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2024)
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2024)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2025)
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2025)
Kinematic Modulation in Driven Spin Resonance
di: Kim, Sunghyun
Pubblicazione: (2026)
di: Kim, Sunghyun
Pubblicazione: (2026)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
di: Wan, Xinyi, et al.
Pubblicazione: (2025)
di: Wan, Xinyi, et al.
Pubblicazione: (2025)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
di: Chen, Shawn Shuoshuo, et al.
Pubblicazione: (2025)
di: Chen, Shawn Shuoshuo, et al.
Pubblicazione: (2025)
POS-ISP: Pipeline Optimization at the Sequence Level for Task-aware ISP
di: Won, Jiyun, et al.
Pubblicazione: (2026)
di: Won, Jiyun, et al.
Pubblicazione: (2026)
Reseña de "Lines in the Sand. Nationalism and identity on the Chilean-Peruvian frontier" de William Skuban
di: Stefanie Gänger
Pubblicazione: (2008)
di: Stefanie Gänger
Pubblicazione: (2008)
WimPyC: an extension module of WimPyDD for the calculation of WIMP capture in celestial bodies
di: Kang, Sunghyun, et al.
Pubblicazione: (2025)
di: Kang, Sunghyun, et al.
Pubblicazione: (2025)
Low-mass constraints on WIMP effective models of inelastic scattering using the Migdal effect
di: Kang, Sunghyun, et al.
Pubblicazione: (2024)
di: Kang, Sunghyun, et al.
Pubblicazione: (2024)
BFS: Back-to-Front Layered Image Synthesis via Knowledge Transfer
di: Kang, Kyoungkook, et al.
Pubblicazione: (2026)
di: Kang, Kyoungkook, et al.
Pubblicazione: (2026)
Leveraging Learned Image Prior for 3D Gaussian Compression
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
Correlation recurrent units: A novel neural architecture for improving the predictive performance of time-series data
di: Sim, Sunghyun, et al.
Pubblicazione: (2022)
di: Sim, Sunghyun, et al.
Pubblicazione: (2022)
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
di: Kim, Dongmin, et al.
Pubblicazione: (2023)
di: Kim, Dongmin, et al.
Pubblicazione: (2023)
EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation
di: Mun, Sunung, et al.
Pubblicazione: (2026)
di: Mun, Sunung, et al.
Pubblicazione: (2026)
VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
di: Kim, Geonung, et al.
Pubblicazione: (2025)
di: Kim, Geonung, et al.
Pubblicazione: (2025)
Locality-aware Gaussian Compression for Fast and High-quality Rendering
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
RNA: Video Editing with ROI-based Neural Atlas
di: Lee, Jaekyeong, et al.
Pubblicazione: (2024)
di: Lee, Jaekyeong, et al.
Pubblicazione: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
di: Miao, Xupeng, et al.
Pubblicazione: (2023)
di: Miao, Xupeng, et al.
Pubblicazione: (2023)
PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
di: Li, Hongpei, et al.
Pubblicazione: (2025)
di: Li, Hongpei, et al.
Pubblicazione: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
di: Zhang, Geng, et al.
Pubblicazione: (2025)
di: Zhang, Geng, et al.
Pubblicazione: (2025)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
di: Jiang, Lijuan, et al.
Pubblicazione: (2025)
di: Jiang, Lijuan, et al.
Pubblicazione: (2025)
Multiscale, Techno-economic Evaluation of Isoreticular Series of CALF-20 for Biogas Upgrading using a Pressure/Vacuum Swing Adsorption (PVSA) Process
di: Shin, Changdon, et al.
Pubblicazione: (2025)
di: Shin, Changdon, et al.
Pubblicazione: (2025)
Relaxation of Projected Prior with Continuous Gap Shrinkage
di: Duan, Leo L, et al.
Pubblicazione: (2026)
di: Duan, Leo L, et al.
Pubblicazione: (2026)
An advance in the arithmetic of the Lie groups as an alternative to the forms of the Campbell-Baker-Hausdorff-Dynkin theorem
di: Kim, Sunghyun, et al.
Pubblicazione: (2024)
di: Kim, Sunghyun, et al.
Pubblicazione: (2024)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
di: Lee, Dong-Jae, et al.
Pubblicazione: (2026)
di: Lee, Dong-Jae, et al.
Pubblicazione: (2026)
A time‐efficient computational binding affinity estimation protocol with utilization of limited experimental data: A case study for adenosine receptor
di: Ilkwon Cho, et al.
Pubblicazione: (2024)
di: Ilkwon Cho, et al.
Pubblicazione: (2024)
IMSE: Intrinsic Mixture of Spectral Experts Fine-tuning for Test-Time Adaptation
di: Baek, Sunghyun, et al.
Pubblicazione: (2026)
di: Baek, Sunghyun, et al.
Pubblicazione: (2026)
The 5 Layer Modernization Stack
di: Aggarwal, Neeraj
Pubblicazione: (2026)
di: Aggarwal, Neeraj
Pubblicazione: (2026)
The Real Cost of Legacy in Digital Payments — And How AI Can Reduce It by 40%
di: Aggarwal, Neeraj
Pubblicazione: (2026)
di: Aggarwal, Neeraj
Pubblicazione: (2026)
Building AI Ready Legacy Systems
di: Aggarwal, Neeraj
Pubblicazione: (2025)
di: Aggarwal, Neeraj
Pubblicazione: (2025)
UICS : A Strategic Framework for Modernizing Tier‑0 Systems
di: Aggarwal, Neeraj
Pubblicazione: (2024)
di: Aggarwal, Neeraj
Pubblicazione: (2024)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
di: Yin, Haofei, et al.
Pubblicazione: (2025)
di: Yin, Haofei, et al.
Pubblicazione: (2025)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
di: Wu, Houming, et al.
Pubblicazione: (2024)
di: Wu, Houming, et al.
Pubblicazione: (2024)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
di: Wang, Hongyu, et al.
Pubblicazione: (2026)
di: Wang, Hongyu, et al.
Pubblicazione: (2026)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
di: Lin, Jun-Liang, et al.
Pubblicazione: (2026)
di: Lin, Jun-Liang, et al.
Pubblicazione: (2026)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUs (Extended Version)
di: Xu, Mingkuan, et al.
Pubblicazione: (2024)
di: Xu, Mingkuan, et al.
Pubblicazione: (2024)
CoherentRaster: Efficient 3D Gaussian Splatting for Light Field Displays
di: Sim, Gyujin, et al.
Pubblicazione: (2026)
di: Sim, Gyujin, et al.
Pubblicazione: (2026)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2024) -
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2025) -
Kinematic Modulation in Driven Spin Resonance
di: Kim, Sunghyun
Pubblicazione: (2026) -
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
di: Wan, Xinyi, et al.
Pubblicazione: (2025) -
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
di: Chen, Shawn Shuoshuo, et al.
Pubblicazione: (2025)