QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Khanna, Danush, Guru, Aditya Kumar, Sridhar, Srivarshinee, Ahmed, Zidan, Bahirwani, Rubhav, Malhotra, Meetu, Jain, Vinija, Chadha, Aman, Das, Amitava, Ghosh, Kripabandhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
by: Chauhan, Anay, et al.
Published: (2026)
by: Chauhan, Anay, et al.
Published: (2026)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
Differentiable Hierarchical Visual Tokenization
by: Aasan, Marius, et al.
Published: (2025)
by: Aasan, Marius, et al.
Published: (2025)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
by: Budagam, Devichand, et al.
Published: (2024)
by: Budagam, Devichand, et al.
Published: (2024)
From Next Token Prediction to (STRIPS) World Models
by: Núñez-Molina, Carlos, et al.
Published: (2025)
by: Núñez-Molina, Carlos, et al.
Published: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
by: Doda, Shravan
Published: (2026)
by: Doda, Shravan
Published: (2026)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
by: Sebők, Miklós, et al.
Published: (2025)
by: Sebők, Miklós, et al.
Published: (2025)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
by: He, Langzhou, et al.
Published: (2026)
by: He, Langzhou, et al.
Published: (2026)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
Hopscotch: Discovering and Skipping Redundancies in Language Models
by: Eyceoz, Mustafa, et al.
Published: (2025)
by: Eyceoz, Mustafa, et al.
Published: (2025)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition
by: Rios, Edwin Arkel, et al.
Published: (2024)
by: Rios, Edwin Arkel, et al.
Published: (2024)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
by: Pitorro, Hugo, et al.
Published: (2025)
by: Pitorro, Hugo, et al.
Published: (2025)
BlankSkip: Early-exit Object Detection onboard Nano-drones
by: Marra, Carlo, et al.
Published: (2026)
by: Marra, Carlo, et al.
Published: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
by: Shen, Meng, et al.
Published: (2026)
by: Shen, Meng, et al.
Published: (2026)
iTrash: Incentivized Token Rewards for Automated Sorting and Handling
by: Ortega, Pablo, et al.
Published: (2025)
by: Ortega, Pablo, et al.
Published: (2025)
Tokenizing Motion: A Generative Approach for Scene Dynamics Compression
by: Yin, Shanzhi, et al.
Published: (2024)
by: Yin, Shanzhi, et al.
Published: (2024)
Next Token Prediction Is a Dead End for Creativity
by: Olatunji, Ibukun, et al.
Published: (2025)
by: Olatunji, Ibukun, et al.
Published: (2025)
Almost Linear Time Consistent Mode Estimation and Quick Shift Clustering
by: Hashemian, Sajjad
Published: (2025)
by: Hashemian, Sajjad
Published: (2025)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
by: Mamidanna, Siddarth, et al.
Published: (2025)
by: Mamidanna, Siddarth, et al.
Published: (2025)
BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers
by: Bhuvaneswaran, Ramshankar, et al.
Published: (2025)
by: Bhuvaneswaran, Ramshankar, et al.
Published: (2025)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Lost in Latent Space: Disentangled Models and the Challenge of Combinatorial Generalisation
by: Montero, Milton L., et al.
Published: (2022)
by: Montero, Milton L., et al.
Published: (2022)
The Battle of LLMs: A Comparative Study in Conversational QA Tasks
by: Rangapur, Aryan, et al.
Published: (2024)
by: Rangapur, Aryan, et al.
Published: (2024)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
by: Lu, Hongyu, et al.
Published: (2026)
by: Lu, Hongyu, et al.
Published: (2026)
VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
by: Kashirskiy, Mark, et al.
Published: (2025)
by: Kashirskiy, Mark, et al.
Published: (2025)
Improving Angular Speed Uniformity by Piecewise Radical Reparameterization
by: Hong, Hoon, et al.
Published: (2024)
by: Hong, Hoon, et al.
Published: (2024)
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
Similar Items
-
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
by: Chauhan, Anay, et al.
Published: (2026) -
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025) -
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025) -
Differentiable Hierarchical Visual Tokenization
by: Aasan, Marius, et al.
Published: (2025) -
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
by: Budagam, Devichand, et al.
Published: (2024)