Draft-based Approximate Inference for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Galim, Kevin, Ewer, Ethan, Kang, Wonjun, Lee, Minjae, Koo, Hyung Il, Lee, Kangwook |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
Parameter-Efficient Fine-Tuning of State Space Models
by: Galim, Kevin, et al.
Published: (2024)
by: Galim, Kevin, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing
by: Kang, Wonjun, et al.
Published: (2024)
by: Kang, Wonjun, et al.
Published: (2024)
Can MLLMs Perform Text-to-Image In-Context Learning?
by: Zeng, Yuchen, et al.
Published: (2024)
by: Zeng, Yuchen, et al.
Published: (2024)
Counting Guidance for High Fidelity Text-to-Image Synthesis
by: Kang, Wonjun, et al.
Published: (2023)
by: Kang, Wonjun, et al.
Published: (2023)
State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
The Expressive Power of Low-Rank Adaptation
by: Zeng, Yuchen, et al.
Published: (2023)
by: Zeng, Yuchen, et al.
Published: (2023)
Prompt-based Depth Pruning of Large Language Models
by: Wee, Juyun, et al.
Published: (2025)
by: Wee, Juyun, et al.
Published: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
by: Lee, Wonjun, et al.
Published: (2026)
by: Lee, Wonjun, et al.
Published: (2026)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking
by: Lee, Jihyun, et al.
Published: (2024)
by: Lee, Jihyun, et al.
Published: (2024)
Prompt-based Learning for Text Readability Assessment
by: Lee, Bruce W., et al.
Published: (2023)
by: Lee, Bruce W., et al.
Published: (2023)
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
by: Lee, Ayeong, et al.
Published: (2025)
by: Lee, Ayeong, et al.
Published: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024)
by: Goel, Raghavv, et al.
Published: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
by: Wen, Zhuofan, et al.
Published: (2024)
by: Wen, Zhuofan, et al.
Published: (2024)
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)
by: Kim, Beomsoo, et al.
Published: (2026)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
by: Cao, Jie, et al.
Published: (2026)
by: Cao, Jie, et al.
Published: (2026)
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
by: Jeon, Sohyeon, et al.
Published: (2025)
by: Jeon, Sohyeon, et al.
Published: (2025)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Program Synthesis via Test-Time Transduction
by: Lee, Kang-il, et al.
Published: (2025)
by: Lee, Kang-il, et al.
Published: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
Personalized LLM Decoding via Contrasting Personal Preference
by: Bu, Hyungjune, et al.
Published: (2025)
by: Bu, Hyungjune, et al.
Published: (2025)
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
Pre-Trained Language Models for Keyphrase Prediction: A Review
by: Umair, Muhammad, et al.
Published: (2024)
by: Umair, Muhammad, et al.
Published: (2024)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
by: Wang, Xuliang, et al.
Published: (2026)
by: Wang, Xuliang, et al.
Published: (2026)
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024)
by: Lad, Vedang, et al.
Published: (2024)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
by: Anshumann, et al.
Published: (2025)
by: Anshumann, et al.
Published: (2025)
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
by: Lee, Bruce W., et al.
Published: (2021)
by: Lee, Bruce W., et al.
Published: (2021)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
How LLMs Might Think
by: Gottlieb, Joseph, et al.
Published: (2026)
by: Gottlieb, Joseph, et al.
Published: (2026)
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
by: Gao, Xiangxiang, et al.
Published: (2024)
by: Gao, Xiangxiang, et al.
Published: (2024)
SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning
by: Verma, Nikhil, et al.
Published: (2025)
by: Verma, Nikhil, et al.
Published: (2025)
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
by: Jeong, Wonjun, et al.
Published: (2025)
by: Jeong, Wonjun, et al.
Published: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Similar Items
-
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
by: Lee, Minjae, et al.
Published: (2026) -
Parameter-Efficient Fine-Tuning of State Space Models
by: Galim, Kevin, et al.
Published: (2024) -
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025) -
Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing
by: Kang, Wonjun, et al.
Published: (2024) -
Can MLLMs Perform Text-to-Image In-Context Learning?
by: Zeng, Yuchen, et al.
Published: (2024)