Similar Items
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
by: Huang, You-Liang, et al.
Published: (2026)
by: Huang, You-Liang, et al.
Published: (2026)
Introducing Support for Move Operations in Melda CRDT
by: Brocco, Amos
Published: (2025)
by: Brocco, Amos
Published: (2025)
Token Level Routing Inference System for Edge Devices
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)
by: Shyam, Vasu, et al.
Published: (2026)
NPB-Rust: NAS Parallel Benchmarks in Rust
by: Martins, Eduardo M., et al.
Published: (2025)
by: Martins, Eduardo M., et al.
Published: (2025)
Extending Contract Verification for Parallel Programming Models to Fortran
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
by: Hinnerskov, Nikolaj Hey, et al.
Published: (2025)
by: Hinnerskov, Nikolaj Hey, et al.
Published: (2025)
Same Engine, Multiple Gears: Parallelizing Fixpoint Iteration at Different Granularities (Extended Version)
by: Kocal, Ali Rasim, et al.
Published: (2026)
by: Kocal, Ali Rasim, et al.
Published: (2026)
ScanWeaver: Compiler-Driven Parallelization of Affine Recurrences via Associative Scan Lowering
by: Wu, Qiying, et al.
Published: (2026)
by: Wu, Qiying, et al.
Published: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
Virtual Garbage Collector (VGC): A Zone-Based Garbage Collection Architecture for Python's Parallel Runtime
by: M, Abdulla
Published: (2025)
by: M, Abdulla
Published: (2025)
Pipeline Parallelism with Controllable Memory
by: Qi, Penghui, et al.
Published: (2024)
by: Qi, Penghui, et al.
Published: (2024)
Mat2Boundary: Treating User-Defined Boundary Condition as SpMV for Distributed PDE Solvers on Block-Structured Grids
by: Cai, Yanzheng, et al.
Published: (2026)
by: Cai, Yanzheng, et al.
Published: (2026)
The State of FaaS: An analysis of public Functions-as-a-Service providers
by: Ekwe-Ekwe, Nnamdi, et al.
Published: (2024)
by: Ekwe-Ekwe, Nnamdi, et al.
Published: (2024)
Regulating Branch Parallelism in LLM Serving
by: Gandhi, Swapnil, et al.
Published: (2026)
by: Gandhi, Swapnil, et al.
Published: (2026)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024)
by: Huang, Mincong, et al.
Published: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Comparing Parallel Functional Array Languages: Programming and Performance
by: van Balen, David, et al.
Published: (2025)
by: van Balen, David, et al.
Published: (2025)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
by: Wang, Yujia, et al.
Published: (2025)
by: Wang, Yujia, et al.
Published: (2025)
Unlocking Full Efficiency of Token Filtering in Large Language Model Training
by: Chai, Di, et al.
Published: (2025)
by: Chai, Di, et al.
Published: (2025)
The Entropy of Parallel Systems
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
Lectures on Parallel Computing
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning
by: Tahir, Anique, et al.
Published: (2024)
by: Tahir, Anique, et al.
Published: (2024)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
by: Cho, Minsik, et al.
Published: (2024)
by: Cho, Minsik, et al.
Published: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Fast and Robust Information Spreading in the Noisy PULL Model
by: D'Archivio, Niccolò, et al.
Published: (2024)
by: D'Archivio, Niccolò, et al.
Published: (2024)
Fermilab's Transition to Token Authentication
by: Dykstra, Dave, et al.
Published: (2025)
by: Dykstra, Dave, et al.
Published: (2025)
FastSet: Parallel Claim Settlement
by: Chen, Xiaohong, et al.
Published: (2025)
by: Chen, Xiaohong, et al.
Published: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
pSTL-Bench: A Micro-Benchmark Suite for Assessing Scalability of C++ Parallel STL Implementations
by: Laso, Ruben, et al.
Published: (2024)
by: Laso, Ruben, et al.
Published: (2024)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Similar Items
-
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026) -
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
by: Huang, You-Liang, et al.
Published: (2026) -
Introducing Support for Move Operations in Melda CRDT
by: Brocco, Amos
Published: (2025) -
Token Level Routing Inference System for Edge Devices
by: She, Jianshu, et al.
Published: (2025) -
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)