Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
Fuente:
arXiv
Saved in:
| Main Authors: | Samragh, Mohammad, Kundu, Arnav, Harrison, David, Nishu, Kumari, Naik, Devang, Cho, Minsik, Farajtabar, Mehrdad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
by: Zibakhsh, Soheil, et al.
Published: (2025)
by: Zibakhsh, Soheil, et al.
Published: (2025)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
by: Nishu, Kumari, et al.
Published: (2025)
by: Nishu, Kumari, et al.
Published: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025)
by: Hannah, Lauren. A, et al.
Published: (2025)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)
by: Nishu, Kumari, et al.
Published: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
by: Cho, Minsik, et al.
Published: (2024)
by: Cho, Minsik, et al.
Published: (2024)
TIDE: Every Layer Knows the Token Beneath the Context
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
by: Bittar, Alexandre, et al.
Published: (2023)
by: Bittar, Alexandre, et al.
Published: (2023)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025)
by: Kim, Han-Byul, et al.
Published: (2025)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024)
by: Alizadeh, Keivan, et al.
Published: (2024)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
by: Hoang, Duc, et al.
Published: (2026)
by: Hoang, Duc, et al.
Published: (2026)
RepCNN: Micro-sized, Mighty Models for Wakeword Detection
by: Kundu, Arnav, et al.
Published: (2024)
by: Kundu, Arnav, et al.
Published: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
An Efficient and Streaming Audio Visual Active Speaker Detection System
by: Kundu, Arnav, et al.
Published: (2024)
by: Kundu, Arnav, et al.
Published: (2024)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
Avey-B
by: Acharya, Devang, et al.
Published: (2026)
by: Acharya, Devang, et al.
Published: (2026)
R2 Loss: Range Restriction Loss for Model Compression and Quantization
by: Kundu, Arnav, et al.
Published: (2023)
by: Kundu, Arnav, et al.
Published: (2023)
Off-diagonally symmetric alternating sign matrices
by: Kumari, Nishu
Published: (2025)
by: Kumari, Nishu
Published: (2025)
A determinantal formula for orthosymplectic Schur functions
by: Kumari, Nishu
Published: (2024)
by: Kumari, Nishu
Published: (2024)
Skew hook Schur functions and the cyclic sieving phenomenon
by: Kumari, Nishu
Published: (2022)
by: Kumari, Nishu
Published: (2022)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
by: Armandpour, Mohammadreza, et al.
Published: (2026)
by: Armandpour, Mohammadreza, et al.
Published: (2026)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
by: Joudaki, Amir, et al.
Published: (2025)
by: Joudaki, Amir, et al.
Published: (2025)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
by: Hoang, Duc N. M, et al.
Published: (2023)
by: Hoang, Duc N. M, et al.
Published: (2023)
NanoKnow: How to Know What Your Language Model Knows
by: Gu, Lingwei, et al.
Published: (2026)
by: Gu, Lingwei, et al.
Published: (2026)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
by: Xiaohu, Xie, et al.
Published: (2026)
by: Xiaohu, Xie, et al.
Published: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Know When To Fold 'Em: Token-Efficient LLM Synthetic Data Generation via Multi-Stage In-Flight Rejection
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
YNTP-100: A Benchmark for Your Next Token Prediction with 100 People
by: Ding, Shiyao, et al.
Published: (2025)
by: Ding, Shiyao, et al.
Published: (2025)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
by: Jovine, Adam S., et al.
Published: (2025)
by: Jovine, Adam S., et al.
Published: (2025)
Similar Items
-
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
by: Zibakhsh, Soheil, et al.
Published: (2025) -
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
by: Nishu, Kumari, et al.
Published: (2025) -
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025) -
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024) -
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)