Saved in:
| Main Authors: | Hahnfeld, Jonas, Blomer, Jakob, Kollegger, Thorsten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.14239 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conceptual Design Report for FAIR Computing
by: Messchendorp, Johan, et al.
Published: (2025)
by: Messchendorp, Johan, et al.
Published: (2025)
Thallus: An RDMA-based Columnar Data Transport Protocol
by: Chakraborty, Jayjeet, et al.
Published: (2024)
by: Chakraborty, Jayjeet, et al.
Published: (2024)
Distributed And Parallel Low-Diameter Decompositions for Arbitrary and Restricted Graphs
by: Dou, Jinfeng, et al.
Published: (2024)
by: Dou, Jinfeng, et al.
Published: (2024)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
by: Latt, Jonas, et al.
Published: (2025)
by: Latt, Jonas, et al.
Published: (2025)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
Can Large Language Models Write Parallel Code?
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
by: Bouvier, Thomas, et al.
Published: (2024)
by: Bouvier, Thomas, et al.
Published: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
BigSUMO: A Scalable Framework for Big Data Traffic Analytics and Parallel Simulation
by: Sengupta, Rahul, et al.
Published: (2026)
by: Sengupta, Rahul, et al.
Published: (2026)
Lectures on Parallel Computing
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
The Entropy of Parallel Systems
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
by: Qing, Yuhao, et al.
Published: (2025)
by: Qing, Yuhao, et al.
Published: (2025)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026)
by: Bu, Tianci, et al.
Published: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
by: Li, Youjia, et al.
Published: (2025)
by: Li, Youjia, et al.
Published: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
by: Lee, Haeun, et al.
Published: (2025)
by: Lee, Haeun, et al.
Published: (2025)
FastSet: Parallel Claim Settlement
by: Chen, Xiaohong, et al.
Published: (2025)
by: Chen, Xiaohong, et al.
Published: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Parallel Seismic Data Processing Performance with Cloud-based Storage
by: Mohapatra, Sasmita, et al.
Published: (2025)
by: Mohapatra, Sasmita, et al.
Published: (2025)
Read-Modify-Writable Snapshots from Read/Write operations
by: Castañeda, Armando, et al.
Published: (2026)
by: Castañeda, Armando, et al.
Published: (2026)
Parallelized Multi-Agent Bayesian Optimization in Lava
by: Snyder, Shay, et al.
Published: (2024)
by: Snyder, Shay, et al.
Published: (2024)
Parallel AIG Refactoring via Conflict Breaking
by: Cai, Ye, et al.
Published: (2024)
by: Cai, Ye, et al.
Published: (2024)
Parallel Gaussian process with kernel approximation in CUDA
by: Carminati, Davide
Published: (2024)
by: Carminati, Davide
Published: (2024)
Adaptive Parallel Downloader for Large Genomic Datasets
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
Dynamic Contract Analysis for Parallel Programming Models
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
Deterministic Parallel High-Quality Hypergraph Partitioning
by: Krause, Robert, et al.
Published: (2025)
by: Krause, Robert, et al.
Published: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
by: Tsolakis, Christos, et al.
Published: (2024)
by: Tsolakis, Christos, et al.
Published: (2024)
FaaS Is Not Enough: Serverless Handling of Burst-Parallel Jobs
by: Barcelona-Pons, Daniel, et al.
Published: (2024)
by: Barcelona-Pons, Daniel, et al.
Published: (2024)
Performance-Driven Optimization of Parallel Breadth-First Search
by: Bhaskar, Marati, et al.
Published: (2025)
by: Bhaskar, Marati, et al.
Published: (2025)
Parallel Spawning Strategies for Dynamic-Aware MPI Applications
by: Martín-Álvarez, Iker, et al.
Published: (2025)
by: Martín-Álvarez, Iker, et al.
Published: (2025)
Parallel Order-Based Core Maintenance in Dynamic Graphs
by: Guo, Bin, et al.
Published: (2022)
by: Guo, Bin, et al.
Published: (2022)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
by: Garner, Kevin, et al.
Published: (2026)
by: Garner, Kevin, et al.
Published: (2026)
HyperParallel: A Supernode-Affinity AI Framework
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Similar Items
-
Conceptual Design Report for FAIR Computing
by: Messchendorp, Johan, et al.
Published: (2025) -
Thallus: An RDMA-based Columnar Data Transport Protocol
by: Chakraborty, Jayjeet, et al.
Published: (2024) -
Distributed And Parallel Low-Diameter Decompositions for Arbitrary and Restricted Graphs
by: Dou, Jinfeng, et al.
Published: (2024) -
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
by: Latt, Jonas, et al.
Published: (2025) -
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)