CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Dong, Yu, Yanxuan, Lengerich, Ben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Plato: Plan to Efficiently Decode for Large Language Model Inference
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Search
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Adaptive Draft-Verification for Efficient Large Language Model Decoding
von: Liu, Xukun, et al.
Veröffentlicht: (2024)
von: Liu, Xukun, et al.
Veröffentlicht: (2024)
AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers
von: Liu, Dong, et al.
Veröffentlicht: (2026)
von: Liu, Dong, et al.
Veröffentlicht: (2026)
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
von: Liu, Dong, et al.
Veröffentlicht: (2026)
von: Liu, Dong, et al.
Veröffentlicht: (2026)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
von: Liu, Dong, et al.
Veröffentlicht: (2026)
von: Liu, Dong, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
Semantic-guided Diverse Decoding for Large Language Model
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG
von: Mo, Fengran, et al.
Veröffentlicht: (2026)
von: Mo, Fengran, et al.
Veröffentlicht: (2026)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
von: Simonds, Toby
Veröffentlicht: (2025)
von: Simonds, Toby
Veröffentlicht: (2025)
DeAL: Decoding-time Alignment for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024)
von: Huang, James Y., et al.
Veröffentlicht: (2024)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
von: Yi, Hanling, et al.
Veröffentlicht: (2024)
von: Yi, Hanling, et al.
Veröffentlicht: (2024)
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
von: Yamaguchi, Atsuki, et al.
Veröffentlicht: (2024)
von: Yamaguchi, Atsuki, et al.
Veröffentlicht: (2024)
Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding
von: Liu, Tianqiao, et al.
Veröffentlicht: (2024)
von: Liu, Tianqiao, et al.
Veröffentlicht: (2024)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Delta -- Contrastive Decoding Mitigates Text Hallucinations in Large Language Models
von: Huang, Cheng Peng, et al.
Veröffentlicht: (2025)
von: Huang, Cheng Peng, et al.
Veröffentlicht: (2025)
Large Language Model as a Universal Clinical Multi-task Decoder
von: Wu, Yujiang, et al.
Veröffentlicht: (2024)
von: Wu, Yujiang, et al.
Veröffentlicht: (2024)
Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
von: Zhao, Zheng, et al.
Veröffentlicht: (2024)
von: Zhao, Zheng, et al.
Veröffentlicht: (2024)
M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
A Survey on Efficient Inference for Large Language Models
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models
von: Wu, Jialiang, et al.
Veröffentlicht: (2025)
von: Wu, Jialiang, et al.
Veröffentlicht: (2025)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models
von: Kim, Seungduk, et al.
Veröffentlicht: (2024)
von: Kim, Seungduk, et al.
Veröffentlicht: (2024)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
von: Gao, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Gao, Xiangxiang, et al.
Veröffentlicht: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Diffusion Language Models Know the Answer Before Decoding
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
Plato: Plan to Efficiently Decode for Large Language Model Inference
von: Jin, Shuowei, et al.
Veröffentlicht: (2024) -
Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Search
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
Adaptive Draft-Verification for Efficient Large Language Model Decoding
von: Liu, Xukun, et al.
Veröffentlicht: (2024)