TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Lijie, Zhang, Zhihao, Chen, Zhuofu, Li, Zikun, Jia, Zhihao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
von: Yao, Feiyu, et al.
Veröffentlicht: (2025)
von: Yao, Feiyu, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
von: Hu, Junhao, et al.
Veröffentlicht: (2026)
von: Hu, Junhao, et al.
Veröffentlicht: (2026)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
RelayLLM: Efficient Reasoning via Collaborative Decoding
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
FastKernels: Benchmarking GPU Kernel Generation in Production
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
Reflection-Window Decoding: Text Generation with Selective Refinement
von: Tang, Zeyu, et al.
Veröffentlicht: (2025)
von: Tang, Zeyu, et al.
Veröffentlicht: (2025)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Controlled Decoding from Language Models
von: Mudgal, Sidharth, et al.
Veröffentlicht: (2023)
von: Mudgal, Sidharth, et al.
Veröffentlicht: (2023)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
von: Bronzini, Marco, et al.
Veröffentlicht: (2025)
von: Bronzini, Marco, et al.
Veröffentlicht: (2025)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
Grammar-Aligned Decoding
von: Park, Kanghee, et al.
Veröffentlicht: (2024)
von: Park, Kanghee, et al.
Veröffentlicht: (2024)
Decoding-based Regression
von: Song, Xingyou, et al.
Veröffentlicht: (2025)
von: Song, Xingyou, et al.
Veröffentlicht: (2025)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2025)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025) -
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025) -
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024) -
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
von: Yao, Feiyu, et al.
Veröffentlicht: (2025) -
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)