Gespeichert in:
| Hauptverfasser: | Liu, Jingyu, Chen, Beidi, Zhang, Ce |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.02789 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)
von: Bachkaniwala, Rajveer, et al.
Veröffentlicht: (2026)
von: Bachkaniwala, Rajveer, et al.
Veröffentlicht: (2026)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Beyond the Speculative Game: A Survey of Speculative Execution in Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
von: Liu, Ziyang
Veröffentlicht: (2026)
von: Liu, Ziyang
Veröffentlicht: (2026)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
von: Sun, Wei, et al.
Veröffentlicht: (2025)
von: Sun, Wei, et al.
Veröffentlicht: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
HAMburger: Accelerating LLM Inference via Token Smashing
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
von: He, Linda, et al.
Veröffentlicht: (2025)
von: He, Linda, et al.
Veröffentlicht: (2025)
Efficient Streaming Language Models with Attention Sinks
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
von: Liu, Yuxin, et al.
Veröffentlicht: (2026)
Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit
von: Goddard, Charles, et al.
Veröffentlicht: (2025)
von: Goddard, Charles, et al.
Veröffentlicht: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
von: Dong, Harry, et al.
Veröffentlicht: (2025)
von: Dong, Harry, et al.
Veröffentlicht: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMs
von: Fu, Yuchen, et al.
Veröffentlicht: (2024)
von: Fu, Yuchen, et al.
Veröffentlicht: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
von: Zhang, Jianguo, et al.
Veröffentlicht: (2025)
von: Zhang, Jianguo, et al.
Veröffentlicht: (2025)
Assessing Large Language Models for Online Extremism Research: Identification, Explanation, and New Knowledge
von: Dong, Beidi, et al.
Veröffentlicht: (2024)
von: Dong, Beidi, et al.
Veröffentlicht: (2024)
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Document-Level In-Context Few-Shot Relation Extraction via Pre-Trained Language Models
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2023)
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2023)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
TiDAR: Think in Diffusion, Talk in Autoregression
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
AlphaToken: Decoupling Adaptation and Stability for Path-Aware Response Token Valuation in LLM Post-Training
von: Qing, Liu, et al.
Veröffentlicht: (2026)
von: Qing, Liu, et al.
Veröffentlicht: (2026)
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)
von: Bachkaniwala, Rajveer, et al.
Veröffentlicht: (2026) -
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026) -
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
von: Fan, Qihang, et al.
Veröffentlicht: (2026) -
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
von: Zhang, Hao, et al.
Veröffentlicht: (2025) -
Beyond the Speculative Game: A Survey of Speculative Execution in Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2024)