Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatia, Nidhi, More, Ankit, Borkar, Ritika, Mitra, Tiyasa, Matas, Ramon, Zhao, Ritchie, Golub, Maximilian, Mudigere, Dheevatsa, Pharris, Brian, Rouhani, Bita Darvish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025)
by: Mitra, Tiyasa, et al.
Published: (2025)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
by: Elango, Venmugil, et al.
Published: (2026)
by: Elango, Venmugil, et al.
Published: (2026)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025)
by: Arfeen, Daiyaan, et al.
Published: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026)
by: Abramovich, Talor, et al.
Published: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
by: Javidnia, Neusha, et al.
Published: (2025)
by: Javidnia, Neusha, et al.
Published: (2025)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
by: Iso, Hayate, et al.
Published: (2026)
by: Iso, Hayate, et al.
Published: (2026)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Enhancing Split Learning with Sharded and Blockchain-Enabled SplitFed Approaches
by: Sokhankhosh, Amirreza, et al.
Published: (2025)
by: Sokhankhosh, Amirreza, et al.
Published: (2025)
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
by: Ai, Mengting, et al.
Published: (2025)
by: Ai, Mengting, et al.
Published: (2025)
Context Parallelism for Scalable Million-Token Inference
by: Yang, Amy, et al.
Published: (2024)
by: Yang, Amy, et al.
Published: (2024)
A Parameterized Decision Tree Classifier to Identify Customers' Online Purchase Intention: An Indian Case
by: Sheel Nidhi Tripathi, et al.
Published: (2025)
by: Sheel Nidhi Tripathi, et al.
Published: (2025)
Emission Distribution for the quantas of Maxwell-Chern-Simon Gauge Field coupled to External Current
by: Kar, Tiyasa
Published: (2021)
by: Kar, Tiyasa
Published: (2021)
‘Never a Colony’?: Rethinking the Colonisation of Enga Province, Papua New Guinea
by: Alex Golub
Published: (2024)
by: Alex Golub
Published: (2024)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
by: Akula, Arjun R., et al.
Published: (2025)
by: Akula, Arjun R., et al.
Published: (2025)
Cascade: Token-Sharded Private LLM Inference
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
Rethinking Risk: Intersectional Inequalities in Long COVID in the United States
by: Bita Nezamdoust
Published: (2026)
by: Bita Nezamdoust
Published: (2026)
ShardTensor: Domain Parallelism for Scientific Machine Learning
by: Adams, Corey, et al.
Published: (2026)
by: Adams, Corey, et al.
Published: (2026)
Introduction to Special Issue ‘Rethinking Decolonisation in Papua New Guinea’
by: Alex Golub, et al.
Published: (2024)
by: Alex Golub, et al.
Published: (2024)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
HelixTrack: Event-Based Tracking and RPM Estimation of Propeller-like Objects
by: Spetlik, Radim, et al.
Published: (2026)
by: Spetlik, Radim, et al.
Published: (2026)
Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
by: Zhang, Ruisi, et al.
Published: (2024)
by: Zhang, Ruisi, et al.
Published: (2024)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025)
by: Wang, Jiangtao, et al.
Published: (2025)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
SP-Chain: Boosting Intra-Shard and Cross-Shard Security and Performance in Blockchain Sharding
by: Li, Mingzhe, et al.
Published: (2024)
by: Li, Mingzhe, et al.
Published: (2024)
DynaShard: Secure and Adaptive Blockchain Sharding Protocol with Hybrid Consensus and Dynamic Shard Management
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Noteworthiness of Sustainable Education in Higher Education: A Qualitative Study
by: Khyati Manchanda, et al.
Published: (2025)
by: Khyati Manchanda, et al.
Published: (2025)
An Investigation into Emotional Intelligence, Foreign Language Anxiety and Empathy through a Cognitive-Affective Course in an EFL Context
by: Ali Rouhani
Published: (2008)
by: Ali Rouhani
Published: (2008)
TRAIL: Cross-Shard Validation for Cryptocurrency Byzantine Shard Protection
by: Jacovetty, Mitch, et al.
Published: (2024)
by: Jacovetty, Mitch, et al.
Published: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
by: Zhang, Geng, et al.
Published: (2025)
by: Zhang, Geng, et al.
Published: (2025)
Drivers and Challenges in Achieving Corporate Carbon Neutrality—Qualitative Investigation of Carbon Capture, Utilization, and Storage Technologies
by: Meena Bhatia, et al.
Published: (2024)
by: Meena Bhatia, et al.
Published: (2024)
Thermodynamic Characteristics of a Fermi Gas with an Invariant Energy Scale and its Astrophysical Implications
by: Kar, Tiyasa, et al.
Published: (2026)
by: Kar, Tiyasa, et al.
Published: (2026)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
by: Qing, Yuhao, et al.
Published: (2025)
by: Qing, Yuhao, et al.
Published: (2025)
Stochastic Approximation with Two Time Scales: The General Case
by: Borkar, Vivek S
Published: (2024)
by: Borkar, Vivek S
Published: (2024)
ShardMemo: Masked MoE Routing for Sharded Agentic LLM Memory
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
Quantitative (Chlorophyll-a) and qualitative (species composition) seasonal fluctuations of phytoplankton in Lavan coastal waters (North of the Persian Gulf)
by: Rouhani Ghadikolaei, K.
Published: (2001)
by: Rouhani Ghadikolaei, K.
Published: (2001)
Influence of Thermostats on the Dynamics of the Helix-Coil Transition
by: Conradi, Maximilian, et al.
Published: (2025)
by: Conradi, Maximilian, et al.
Published: (2025)
Nonequilibrium Dynamics of the Helix-Coil Transition in Polyalanine
by: Conradi, Maximilian, et al.
Published: (2025)
by: Conradi, Maximilian, et al.
Published: (2025)
Similar Items
-
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025) -
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
by: Elango, Venmugil, et al.
Published: (2026) -
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025) -
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026) -
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)