TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Xiaojuan, Meng, Fanxu, Tang, Pingzhi, Wang, Yuxuan, Yin, Di, Sun, Xing, Zhang, Muhan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TransMLA: Multi-Head Latent Attention Is All You Need
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning
von: Meng, Fanxu, et al.
Veröffentlicht: (2024)
von: Meng, Fanxu, et al.
Veröffentlicht: (2024)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
von: Meng, Fanxu
Veröffentlicht: (2026)
von: Meng, Fanxu
Veröffentlicht: (2026)
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
von: Xu, Yufei, et al.
Veröffentlicht: (2026)
von: Xu, Yufei, et al.
Veröffentlicht: (2026)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
von: Meng, Fanxu, et al.
Veröffentlicht: (2024)
von: Meng, Fanxu, et al.
Veröffentlicht: (2024)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
von: Li, Yanan, et al.
Veröffentlicht: (2025)
von: Li, Yanan, et al.
Veröffentlicht: (2025)
RulE: Knowledge Graph Reasoning with Rule Embedding
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2022)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2022)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
Parrot Mind: Towards Explaining the Complex Task Reasoning of Pretrained Large Language Models with Template-Content Structure
von: Yang, Haotong, et al.
Veröffentlicht: (2023)
von: Yang, Haotong, et al.
Veröffentlicht: (2023)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
von: Zhang, Hengrui, et al.
Veröffentlicht: (2025)
von: Zhang, Hengrui, et al.
Veröffentlicht: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
von: Wang, Meixuan, et al.
Veröffentlicht: (2025)
von: Wang, Meixuan, et al.
Veröffentlicht: (2025)
Mars: Situated Inductive Reasoning in an Open-World Environment
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2024)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2024)
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
von: Song, Mingcong, et al.
Veröffentlicht: (2024)
von: Song, Mingcong, et al.
Veröffentlicht: (2024)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
von: Liu, Ziyang
Veröffentlicht: (2026)
von: Liu, Ziyang
Veröffentlicht: (2026)
Towards Low-bit Communication for Tensor Parallel LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
Communication Compression for Tensor Parallel LLM Inference
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
Law in Silico: Simulating Legal Society with LLM-Based Agents
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
Case-Based or Rule-Based: How Do Transformers Do the Math?
von: Hu, Yi, et al.
Veröffentlicht: (2024)
von: Hu, Yi, et al.
Veröffentlicht: (2024)
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
von: Guanzhong, Chen
Veröffentlicht: (2026)
von: Guanzhong, Chen
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
von: Li, Luchang, et al.
Veröffentlicht: (2026)
von: Li, Luchang, et al.
Veröffentlicht: (2026)
Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility
von: Oh, Jungsuk, et al.
Veröffentlicht: (2026)
von: Oh, Jungsuk, et al.
Veröffentlicht: (2026)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-Decoding
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
von: An, Tai, et al.
Veröffentlicht: (2025)
von: An, Tai, et al.
Veröffentlicht: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
von: Song, Chendong, et al.
Veröffentlicht: (2026)
von: Song, Chendong, et al.
Veröffentlicht: (2026)
LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning
von: Mao, Yansheng, et al.
Veröffentlicht: (2024)
von: Mao, Yansheng, et al.
Veröffentlicht: (2024)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TransMLA: Multi-Head Latent Attention Is All You Need
von: Meng, Fanxu, et al.
Veröffentlicht: (2025) -
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning
von: Meng, Fanxu, et al.
Veröffentlicht: (2024) -
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026) -
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
von: Meng, Fanxu
Veröffentlicht: (2026) -
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
von: Xu, Yufei, et al.
Veröffentlicht: (2026)