How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Minghua, Zhang, Lingzhe, Liu, Yuan, Zhou, Xiao, Liu, Aiwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
von: Shen, Siyuan, et al.
Veröffentlicht: (2025)
von: Shen, Siyuan, et al.
Veröffentlicht: (2025)
Spatiotemporal Analysis of Parallelized Computing at the Extreme Edge
von: Nabil, Yasser, et al.
Veröffentlicht: (2025)
von: Nabil, Yasser, et al.
Veröffentlicht: (2025)
Robust Recursive Query Parallelism in Graph Database Management Systems
von: Chakraborty, Anurag, et al.
Veröffentlicht: (2025)
von: Chakraborty, Anurag, et al.
Veröffentlicht: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
PEVLM: Parallel Encoding for Vision-Language Models
von: Kang, Letian, et al.
Veröffentlicht: (2025)
von: Kang, Letian, et al.
Veröffentlicht: (2025)
GPU-Accelerated Parallel Selected Inversion for Structured Matrices Using sTiles
von: Fattah, Esmail Abdul, et al.
Veröffentlicht: (2025)
von: Fattah, Esmail Abdul, et al.
Veröffentlicht: (2025)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
von: Lazcano, Raquel, et al.
Veröffentlicht: (2024)
von: Lazcano, Raquel, et al.
Veröffentlicht: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
von: Lin, Wei-Fen, et al.
Veröffentlicht: (2026)
von: Lin, Wei-Fen, et al.
Veröffentlicht: (2026)
Comparing Parallel Functional Array Languages: Programming and Performance
von: van Balen, David, et al.
Veröffentlicht: (2025)
von: van Balen, David, et al.
Veröffentlicht: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2024)
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2024)
Kino-PAX: Highly Parallel Kinodynamic Sampling-based Planner
von: Perrault, Nicolas, et al.
Veröffentlicht: (2024)
von: Perrault, Nicolas, et al.
Veröffentlicht: (2024)
Parallel $k$d-tree with Batch Updates
von: Men, Ziyang, et al.
Veröffentlicht: (2024)
von: Men, Ziyang, et al.
Veröffentlicht: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Can Large Language Models Predict Parallel Code Performance?
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Using UML State Diagrams for Modelling the Performance of Parallel Programs
von: Jorge Ortega Arjona
Veröffentlicht: (2008)
von: Jorge Ortega Arjona
Veröffentlicht: (2008)
CPMA: An Efficient Batch-Parallel Compressed Set Without Pointers
von: Wheatman, Brian, et al.
Veröffentlicht: (2023)
von: Wheatman, Brian, et al.
Veröffentlicht: (2023)
PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
von: Yang, Liu, et al.
Veröffentlicht: (2026)
von: Yang, Liu, et al.
Veröffentlicht: (2026)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025) -
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
von: Huang, Haochen, et al.
Veröffentlicht: (2025) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025) -
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
von: Shen, Siyuan, et al.
Veröffentlicht: (2025) -
Spatiotemporal Analysis of Parallelized Computing at the Extreme Edge
von: Nabil, Yasser, et al.
Veröffentlicht: (2025)