TokenButler: Token Importance is Predictable
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akhauri, Yash, AbouElhamayed, Ahmed F, Gao, Yifei, Chang, Chi-Chih, Gobriel, Sameh, Jain, Nilesh, Abdelfattah, Mohamed S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
SplitReason: Learning To Offload Reasoning
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
Attamba: Attending To Multi-Token States
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2024)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2024)
Compute Where it Counts: Self Optimizing Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2026)
von: Akhauri, Yash, et al.
Veröffentlicht: (2026)
Encodings for Prediction-based Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
Regression Language Models for Code
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
On Latency Predictors for Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2023)
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2023)
Learning to Reason with Mixture of Tokens
von: Jain, Adit, et al.
Veröffentlicht: (2025)
von: Jain, Adit, et al.
Veröffentlicht: (2025)
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
von: Hu, Zhanqiu, et al.
Veröffentlicht: (2025)
von: Hu, Zhanqiu, et al.
Veröffentlicht: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
von: Singh, Janvijay, et al.
Veröffentlicht: (2026)
NITRO: LLM Inference on Intel Laptop NPUs
von: Fei, Anthony, et al.
Veröffentlicht: (2024)
von: Fei, Anthony, et al.
Veröffentlicht: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
TECP: Token-Entropy Conformal Prediction for LLMs
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
KV Prediction for Improved Time to First Token
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
The Token Tax: Systematic Bias in Multilingual Tokenization
von: Lundin, Jessica M., et al.
Veröffentlicht: (2025)
von: Lundin, Jessica M., et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
State over Tokens: Characterizing the Role of Reasoning Tokens
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy
von: Ahmed, Amr
Veröffentlicht: (2026)
von: Ahmed, Amr
Veröffentlicht: (2026)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
Efficient Joint Prediction of Multiple Future Tokens
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2024)
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2024)
Alternatives To Next Token Prediction In Text Generation -- A Survey
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
von: Zhang, Zihou, et al.
Veröffentlicht: (2026)
von: Zhang, Zihou, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025) -
SplitReason: Learning To Offload Reasoning
von: Akhauri, Yash, et al.
Veröffentlicht: (2025) -
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2024) -
Attamba: Attending To Multi-Token States
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)