Saved in:
| Main Authors: | Los, Denis, Petushkov, Igor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.01222 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
by: Los, Denis, et al.
Published: (2025)
by: Los, Denis, et al.
Published: (2025)
Dynamic Simultaneous Multithreaded Architecture
by: Ortiz-Arroyo, Daniel, et al.
Published: (2024)
by: Ortiz-Arroyo, Daniel, et al.
Published: (2024)
Multithreaded Fine-Grained Asynchronous BSP for Integer Sorting with LCI and OpenMP
by: Cheng, Minyu, et al.
Published: (2026)
by: Cheng, Minyu, et al.
Published: (2026)
Examining MPI and its Extensions for Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
LCI: a Lightweight Communication Interface for Efficient Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Multithreaded parallelism for heterogeneous clusters of QPUs
by: Seitz, Philipp, et al.
Published: (2023)
by: Seitz, Philipp, et al.
Published: (2023)
LB4OMP: A Dynamic Load Balancing Library for Multithreaded Applications
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
by: Li, Cong, et al.
Published: (2025)
by: Li, Cong, et al.
Published: (2025)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Parallel Order-Based Core Maintenance in Dynamic Graphs
by: Guo, Bin, et al.
Published: (2022)
by: Guo, Bin, et al.
Published: (2022)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
by: Tsolakis, Christos, et al.
Published: (2024)
by: Tsolakis, Christos, et al.
Published: (2024)
Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation
by: Zhang, Zikai, et al.
Published: (2026)
by: Zhang, Zikai, et al.
Published: (2026)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Fine-grained MoE Load Balancing with Linear Programming
by: Zhao, Chenqi, et al.
Published: (2025)
by: Zhao, Chenqi, et al.
Published: (2025)
Efficient Task Graph Scheduling for Parallel QR Factorization in SLSQP
by: Chatterjee, Soumyajit, et al.
Published: (2025)
by: Chatterjee, Soumyajit, et al.
Published: (2025)
A Parallel and Distributed Rust Library for Core Decomposition on Large Graphs
by: Rucci, Davide, et al.
Published: (2025)
by: Rucci, Davide, et al.
Published: (2025)
Parallel $k$-Core Decomposition with Batched Updates and Asynchronous Reads
by: Liu, Quanquan C., et al.
Published: (2024)
by: Liu, Quanquan C., et al.
Published: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
by: Schieffer, Gabin, et al.
Published: (2026)
by: Schieffer, Gabin, et al.
Published: (2026)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Linear Complexity $\mathcal{H}^2$ Direct Solver for Fine-Grained Parallel Architectures
by: Boukaram, Wajih, et al.
Published: (2025)
by: Boukaram, Wajih, et al.
Published: (2025)
Deferred Objects to Enhance Smart Contract Programming with Optimistic Parallel Execution
by: Mitenkov, George, et al.
Published: (2024)
by: Mitenkov, George, et al.
Published: (2024)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
by: Dutt, Anurag, et al.
Published: (2025)
by: Dutt, Anurag, et al.
Published: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
VersaSlot: Efficient Fine-grained FPGA Sharing with Big.Little Slots and Live Migration in FPGA Cluster
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
by: Maeda, Yuki, et al.
Published: (2026)
by: Maeda, Yuki, et al.
Published: (2026)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Lectures on Parallel Computing
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
The Entropy of Parallel Systems
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
by: Idouar, Yacine, et al.
Published: (2025)
by: Idouar, Yacine, et al.
Published: (2025)
Similar Items
-
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
by: Los, Denis, et al.
Published: (2025) -
Dynamic Simultaneous Multithreaded Architecture
by: Ortiz-Arroyo, Daniel, et al.
Published: (2024) -
Multithreaded Fine-Grained Asynchronous BSP for Integer Sorting with LCI and OpenMP
by: Cheng, Minyu, et al.
Published: (2026) -
Examining MPI and its Extensions for Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025) -
LCI: a Lightweight Communication Interface for Efficient Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025)