Communication Compression for Tensor Parallel LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hansen-Palmus, Jan, Le, Michael Truong, Hausdörfer, Oliver, Verma, Alok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LatentLLM: Attention-Aware Joint Tensor Compression
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Towards Low-bit Communication for Tensor Parallel LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
von: Adeseye, Aisvarya, et al.
Veröffentlicht: (2026)
von: Adeseye, Aisvarya, et al.
Veröffentlicht: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
Transparent Screening for LLM Inference and Training Impacts
von: Pachot, Arnault, et al.
Veröffentlicht: (2026)
von: Pachot, Arnault, et al.
Veröffentlicht: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
von: Yi, Hanling, et al.
Veröffentlicht: (2024)
von: Yi, Hanling, et al.
Veröffentlicht: (2024)
Set-LLM: A Permutation-Invariant LLM
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
von: Zhong, Tianle, et al.
Veröffentlicht: (2026)
von: Zhong, Tianle, et al.
Veröffentlicht: (2026)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
von: Gokden, Burc
Veröffentlicht: (2025)
von: Gokden, Burc
Veröffentlicht: (2025)
PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database
von: Sun, Hui, et al.
Veröffentlicht: (2025)
von: Sun, Hui, et al.
Veröffentlicht: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
Vidur: A Large-Scale Simulation Framework For LLM Inference
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
KV Cache Transform Coding for Compact Storage in LLM Inference
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Locally Coherent Parallel Decoding in Diffusion Language Models
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
von: Chen, Guoxuan, et al.
Veröffentlicht: (2024)
von: Chen, Guoxuan, et al.
Veröffentlicht: (2024)
Exploring and Improving Drafts in Blockwise Parallel Decoding
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
von: Alipour, Mohammadsajad, et al.
Veröffentlicht: (2025)
von: Alipour, Mohammadsajad, et al.
Veröffentlicht: (2025)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity
von: Bae, Henry, et al.
Veröffentlicht: (2023)
von: Bae, Henry, et al.
Veröffentlicht: (2023)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024)
von: He, Junhui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LatentLLM: Attention-Aware Joint Tensor Compression
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025) -
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025) -
Towards Low-bit Communication for Tensor Parallel LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024) -
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)