TokenButler: Token Importance is Predictable
Fuente:
arXiv
Salvato in:
| Autori principali: | Akhauri, Yash, AbouElhamayed, Ahmed F, Gao, Yifei, Chang, Chi-Chih, Gobriel, Sameh, Jain, Nilesh, Abdelfattah, Mohamed S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
SplitReason: Learning To Offload Reasoning
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
di: Dotzel, Jordan, et al.
Pubblicazione: (2024)
di: Dotzel, Jordan, et al.
Pubblicazione: (2024)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
Attamba: Attending To Multi-Token States
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2024)
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2024)
Compute Where it Counts: Self Optimizing Language Models
di: Akhauri, Yash, et al.
Pubblicazione: (2026)
di: Akhauri, Yash, et al.
Pubblicazione: (2026)
Encodings for Prediction-based Neural Architecture Search
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
Regression Language Models for Code
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
On Latency Predictors for Neural Architecture Search
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2023)
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2023)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2023)
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2023)
Learning to Reason with Mixture of Tokens
di: Jain, Adit, et al.
Pubblicazione: (2025)
di: Jain, Adit, et al.
Pubblicazione: (2025)
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
di: Hu, Zhanqiu, et al.
Pubblicazione: (2025)
di: Hu, Zhanqiu, et al.
Pubblicazione: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
NITRO: LLM Inference on Intel Laptop NPUs
di: Fei, Anthony, et al.
Pubblicazione: (2024)
di: Fei, Anthony, et al.
Pubblicazione: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
di: Liu, Peijie, et al.
Pubblicazione: (2025)
di: Liu, Peijie, et al.
Pubblicazione: (2025)
Cautious Next Token Prediction
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
TECP: Token-Entropy Conformal Prediction for LLMs
di: Xu, Beining, et al.
Pubblicazione: (2025)
di: Xu, Beining, et al.
Pubblicazione: (2025)
KV Prediction for Improved Time to First Token
di: Horton, Maxwell, et al.
Pubblicazione: (2024)
di: Horton, Maxwell, et al.
Pubblicazione: (2024)
The Token Tax: Systematic Bias in Multilingual Tokenization
di: Lundin, Jessica M., et al.
Pubblicazione: (2025)
di: Lundin, Jessica M., et al.
Pubblicazione: (2025)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
State over Tokens: Characterizing the Role of Reasoning Tokens
di: Levy, Mosh, et al.
Pubblicazione: (2025)
di: Levy, Mosh, et al.
Pubblicazione: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy
di: Ahmed, Amr
Pubblicazione: (2026)
di: Ahmed, Amr
Pubblicazione: (2026)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
Efficient Joint Prediction of Multiple Future Tokens
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
Text Generation Beyond Discrete Token Sampling
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
di: Goel, Vatsal, et al.
Pubblicazione: (2026)
di: Goel, Vatsal, et al.
Pubblicazione: (2026)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
di: Wang, Dixuan, et al.
Pubblicazione: (2024)
di: Wang, Dixuan, et al.
Pubblicazione: (2024)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2024)
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2024)
Alternatives To Next Token Prediction In Text Generation -- A Survey
di: Wyatt, Charlie, et al.
Pubblicazione: (2025)
di: Wyatt, Charlie, et al.
Pubblicazione: (2025)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
di: Aynetdinov, Ansar, et al.
Pubblicazione: (2025)
di: Aynetdinov, Ansar, et al.
Pubblicazione: (2025)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
di: Zhang, Zihou, et al.
Pubblicazione: (2026)
di: Zhang, Zihou, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025) -
SplitReason: Learning To Offload Reasoning
di: Akhauri, Yash, et al.
Pubblicazione: (2025) -
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
di: Dotzel, Jordan, et al.
Pubblicazione: (2024) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
di: Akhauri, Yash, et al.
Pubblicazione: (2024) -
Attamba: Attending To Multi-Token States
di: Akhauri, Yash, et al.
Pubblicazione: (2024)