Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
Fuente:
arXiv
Saved in:
| Main Authors: | Paliotta, Daniele, Wang, Junxiong, Pagliardini, Matteo, Li, Kevin Y., Bick, Aviv, Kolter, J. Zico, Gu, Albert, Fleuret, François, Dao, Tri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
Mamba-3: Improved Sequence Modeling using State Space Principles
by: Lahoti, Aakash, et al.
Published: (2026)
by: Lahoti, Aakash, et al.
Published: (2026)
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)
by: González, Ramón Calvo, et al.
Published: (2025)
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
by: Wang, Junxiong, et al.
Published: (2025)
by: Wang, Junxiong, et al.
Published: (2025)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026)
by: Bick, Aviv, et al.
Published: (2026)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
by: Ai, Xinyue, et al.
Published: (2025)
by: Ai, Xinyue, et al.
Published: (2025)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
by: Huang, Benhao, et al.
Published: (2026)
by: Huang, Benhao, et al.
Published: (2026)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
by: Gu, Albert, et al.
Published: (2023)
by: Gu, Albert, et al.
Published: (2023)
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
by: Dao, Tri, et al.
Published: (2024)
by: Dao, Tri, et al.
Published: (2024)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
by: Kim, Eungyeup, et al.
Published: (2026)
by: Kim, Eungyeup, et al.
Published: (2026)
One-Step Diffusion Distillation through Score Implicit Matching
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
by: Christakopoulou, Konstantina, et al.
Published: (2024)
by: Christakopoulou, Konstantina, et al.
Published: (2024)
TwiSTAR:Think Fast, Think Slow, Then Act,Generative Recommendation with Adaptive Reasoning
by: Cao, Shiteng, et al.
Published: (2026)
by: Cao, Shiteng, et al.
Published: (2026)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
by: Hwang, Sukjun, et al.
Published: (2024)
by: Hwang, Sukjun, et al.
Published: (2024)
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
by: Su, DiJia, et al.
Published: (2024)
by: Su, DiJia, et al.
Published: (2024)
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
by: Zhang, Junyu, et al.
Published: (2025)
by: Zhang, Junyu, et al.
Published: (2025)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
by: Duan, Xintong, et al.
Published: (2025)
by: Duan, Xintong, et al.
Published: (2025)
Thinker: Learning to Think Fast and Slow
by: Chung, Stephen, et al.
Published: (2025)
by: Chung, Stephen, et al.
Published: (2025)
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
by: Pan, Qianjun, et al.
Published: (2025)
by: Pan, Qianjun, et al.
Published: (2025)
Similar Items
-
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024) -
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
by: Wang, Junxiong, et al.
Published: (2024) -
Mamba-3: Improved Sequence Modeling using State Space Principles
by: Lahoti, Aakash, et al.
Published: (2026) -
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025) -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
by: Wang, Junxiong, et al.
Published: (2025)