Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
Fuente:
arXiv
Saved in:
| Main Authors: | Deschenaux, Justin, Gulcehre, Caglar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
by: Deschenaux, Justin, et al.
Published: (2026)
by: Deschenaux, Justin, et al.
Published: (2026)
Promises, Outlooks and Challenges of Diffusion Language Modeling
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
Partition Generative Modeling: Masked Modeling Without Masks
by: Deschenaux, Justin, et al.
Published: (2025)
by: Deschenaux, Justin, et al.
Published: (2025)
The Diffusion Duality, Chapter II: $Ψ$-Samplers
by: Deschenaux, Justin, et al.
Published: (2026)
by: Deschenaux, Justin, et al.
Published: (2026)
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
by: Jo, Mingyu, et al.
Published: (2025)
by: Jo, Mingyu, et al.
Published: (2025)
Self-Recognition in Language Models
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
Scaling Beyond Masked Diffusion Language Models
by: Sahoo, Subham Sekhar, et al.
Published: (2026)
by: Sahoo, Subham Sekhar, et al.
Published: (2026)
Efficient Knowledge Injection in LLMs via Self-Distillation
by: Kujanpää, Kalle, et al.
Published: (2024)
by: Kujanpää, Kalle, et al.
Published: (2024)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
The Diffusion Duality
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
by: Ma, Liqun, et al.
Published: (2024)
by: Ma, Liqun, et al.
Published: (2024)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
by: Klein, Lars, et al.
Published: (2024)
by: Klein, Lars, et al.
Published: (2024)
Self-Calibrating Language Models via Test-Time Discriminative Distillation
by: Hedna, Mohamed Rissal, et al.
Published: (2026)
by: Hedna, Mohamed Rissal, et al.
Published: (2026)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Distilled Self-Critique of LLMs with Synthetic Data: a Bayesian Perspective
by: Gallego, Victor
Published: (2023)
by: Gallego, Victor
Published: (2023)
AutoTimes: Autoregressive Time Series Forecasters via Large Language Models
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026)
by: Kirchenbauer, John, et al.
Published: (2026)
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
by: Ye, Jiacheng, et al.
Published: (2024)
by: Ye, Jiacheng, et al.
Published: (2024)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
by: Terekhov, Mikhail, et al.
Published: (2024)
by: Terekhov, Mikhail, et al.
Published: (2024)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Self-Refining Language Model Anonymizers via Adversarial Distillation
by: Kim, Kyuyoung, et al.
Published: (2025)
by: Kim, Kyuyoung, et al.
Published: (2025)
Few-Step Diffusion Language Models via Trajectory Self-Distillation
by: Zhang, Tunyu, et al.
Published: (2026)
by: Zhang, Tunyu, et al.
Published: (2026)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
by: Hoshino, Yuichiro, et al.
Published: (2025)
by: Hoshino, Yuichiro, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
by: Ma, Ruotian, et al.
Published: (2025)
by: Ma, Ruotian, et al.
Published: (2025)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
by: Kiruluta, Andrew
Published: (2026)
by: Kiruluta, Andrew
Published: (2026)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
by: Matrenok, Simon, et al.
Published: (2025)
by: Matrenok, Simon, et al.
Published: (2025)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
Similar Items
-
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
by: Deschenaux, Justin, et al.
Published: (2026) -
Promises, Outlooks and Challenges of Diffusion Language Modeling
by: Deschenaux, Justin, et al.
Published: (2024) -
Partition Generative Modeling: Masked Modeling Without Masks
by: Deschenaux, Justin, et al.
Published: (2025) -
The Diffusion Duality, Chapter II: $Ψ$-Samplers
by: Deschenaux, Justin, et al.
Published: (2026) -
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
by: Jo, Mingyu, et al.
Published: (2025)