Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bergner, Benjamin, Skliar, Andrii, Royer, Amelie, Blankevoort, Tijmen, Asano, Yuki, Bejnordi, Babak Ehteshami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
von: Bejnordi, Babak Ehteshami, et al.
Veröffentlicht: (2024)
von: Bejnordi, Babak Ehteshami, et al.
Veröffentlicht: (2024)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
VeRA: Vector-based Random Matrix Adaptation
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2023)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2023)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2026)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2026)
Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
SpinQuant: LLM quantization with learned rotations
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
The LLM Surgeon
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2023)
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2023)
Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning
von: Tang, Lexiang, et al.
Veröffentlicht: (2026)
von: Tang, Lexiang, et al.
Veröffentlicht: (2026)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
von: He, Yanjie
Veröffentlicht: (2026)
von: He, Yanjie
Veröffentlicht: (2026)
TiDAR: Think in Diffusion, Talk in Autoregression
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction
von: Glavas, Theodore, et al.
Veröffentlicht: (2026)
von: Glavas, Theodore, et al.
Veröffentlicht: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
von: Huang, Hui, et al.
Veröffentlicht: (2025)
von: Huang, Hui, et al.
Veröffentlicht: (2025)
QuickMerge++: Fast Token Merging with Autoregressive Prior
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
von: Hu, Mengya, et al.
Veröffentlicht: (2024)
von: Hu, Mengya, et al.
Veröffentlicht: (2024)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
von: He, Jiashu, et al.
Veröffentlicht: (2026)
von: He, Jiashu, et al.
Veröffentlicht: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
von: Barron, Joshua, et al.
Veröffentlicht: (2025)
von: Barron, Joshua, et al.
Veröffentlicht: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
SLM as Guardian: Pioneering AI Safety with Small Language Models
von: Kwon, Ohjoon, et al.
Veröffentlicht: (2024)
von: Kwon, Ohjoon, et al.
Veröffentlicht: (2024)
Kuwain 1.5B: An Arabic SLM via Language Injection
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
Edit-Constrained Decoding for Sentence Simplification
von: Zetsu, Tatsuya, et al.
Veröffentlicht: (2024)
von: Zetsu, Tatsuya, et al.
Veröffentlicht: (2024)
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
von: Du, Y., et al.
Veröffentlicht: (2025)
von: Du, Y., et al.
Veröffentlicht: (2025)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
Steering LLM Thinking with Budget Guidance
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
Thinkless: LLM Learns When to Think
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
von: Laitenberger, Filipe, et al.
Veröffentlicht: (2025)
von: Laitenberger, Filipe, et al.
Veröffentlicht: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
Uncovering Autoregressive LLM Knowledge of Thematic Fit in Event Representation
von: Alshemali, Safeyah Khaled, et al.
Veröffentlicht: (2024)
von: Alshemali, Safeyah Khaled, et al.
Veröffentlicht: (2024)
Thinking Before Constraining: A Unified Decoding Framework for Large Language Models
von: Nguyen, Ngoc Trinh Hung, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc Trinh Hung, et al.
Veröffentlicht: (2026)
Fast Autoregressive Video Generation with Diagonal Decoding
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
von: Dudeja, Divij, et al.
Veröffentlicht: (2025)
von: Dudeja, Divij, et al.
Veröffentlicht: (2025)
Fast-Slow-Thinking: Complex Task Solving with Large Language Models
von: Sun, Yiliu, et al.
Veröffentlicht: (2025)
von: Sun, Yiliu, et al.
Veröffentlicht: (2025)
Think While You Write: Hypothesis Verification Promotes Faithful Knowledge-to-Text Generation
von: Qiu, Yifu, et al.
Veröffentlicht: (2023)
von: Qiu, Yifu, et al.
Veröffentlicht: (2023)
MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
von: Bejnordi, Babak Ehteshami, et al.
Veröffentlicht: (2024) -
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024) -
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024) -
VeRA: Vector-based Random Matrix Adaptation
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2023) -
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2026)