Learning Adaptive LLM Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Chloe H., Ye, Zhe, Tenka, Samuel, Yang, Aidan, Kong, Soonho, Ghai, Udaya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intent-aligned Formal Specification Synthesis via Traceable Refinement
by: Ye, Zhe, et al.
Published: (2026)
by: Ye, Zhe, et al.
Published: (2026)
Sample-Optimal Agnostic Boosting with Unlabeled Data
by: Ghai, Udaya, et al.
Published: (2025)
by: Ghai, Udaya, et al.
Published: (2025)
Sample-Efficient Agnostic Boosting
by: Ghai, Udaya, et al.
Published: (2024)
by: Ghai, Udaya, et al.
Published: (2024)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026)
by: Flamant, Cedric, et al.
Published: (2026)
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026)
by: Deb, Rohan, et al.
Published: (2026)
Outbound Modeling for Inventory Management
by: Savorgnan, Riccardo, et al.
Published: (2025)
by: Savorgnan, Riccardo, et al.
Published: (2025)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Neural Coordination and Capacity Control for Inventory Management
by: Eisenach, Carson, et al.
Published: (2024)
by: Eisenach, Carson, et al.
Published: (2024)
Adaptive Policy Learning Under Unknown Network Interference
by: Gleich, Aidan, et al.
Published: (2026)
by: Gleich, Aidan, et al.
Published: (2026)
Residual Learning and Context Encoding for Adaptive Offline-to-Online Reinforcement Learning
by: Nakhaei, Mohammadreza, et al.
Published: (2024)
by: Nakhaei, Mohammadreza, et al.
Published: (2024)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Teaching LLMs Program Semantics via Symbolic Execution Traces
by: Bayer, Jonas, et al.
Published: (2026)
by: Bayer, Jonas, et al.
Published: (2026)
Interactive Diabetes Risk Prediction Using Explainable Machine Learning: A Dash-Based Approach with SHAP, LIME, and Comorbidity Insights
by: Allani, Udaya
Published: (2025)
by: Allani, Udaya
Published: (2025)
MADQRL: Distributed Quantum Reinforcement Learning Framework for Multi-Agent Environments
by: Sawaika, Abhishek, et al.
Published: (2026)
by: Sawaika, Abhishek, et al.
Published: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
by: Gao, Lei, et al.
Published: (2025)
by: Gao, Lei, et al.
Published: (2025)
Few-Shot Concept Unlearning with Low Rank Adaptation
by: Shreyas, Udaya, et al.
Published: (2025)
by: Shreyas, Udaya, et al.
Published: (2025)
Adaptive Budget Allocation in LLM-Augmented Surveys
by: Ye, Zikun, et al.
Published: (2026)
by: Ye, Zikun, et al.
Published: (2026)
Stroke Prediction using Clinical and Social Features in Machine Learning
by: Chadha, Aidan
Published: (2024)
by: Chadha, Aidan
Published: (2024)
Learning from Similar Linear Representations: Adaptivity, Minimaxity, and Robustness
by: Tian, Ye, et al.
Published: (2023)
by: Tian, Ye, et al.
Published: (2023)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
by: Lu, Jialin, et al.
Published: (2026)
by: Lu, Jialin, et al.
Published: (2026)
Tokenized Bandit for LLM Decoding and Alignment
by: Shin, Suho, et al.
Published: (2025)
by: Shin, Suho, et al.
Published: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
by: Xu, WeiZhe, et al.
Published: (2026)
by: Xu, WeiZhe, et al.
Published: (2026)
Entropy Regularized Task Representation Learning for Offline Meta-Reinforcement Learning
by: Nakhaei, Mohammadreza, et al.
Published: (2024)
by: Nakhaei, Mohammadreza, et al.
Published: (2024)
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
by: Ye, Zicong, et al.
Published: (2025)
by: Ye, Zicong, et al.
Published: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
ElfCore: A 28nm Neural Processor Enabling Dynamic Structured Sparse Training and Online Self-Supervised Learning with Activity-Dependent Weight Update
by: Su, Zhe, et al.
Published: (2025)
by: Su, Zhe, et al.
Published: (2025)
BRIDGE: Building Representations In Domain Guided Program Synthesis
by: George, Robert Joseph, et al.
Published: (2025)
by: George, Robert Joseph, et al.
Published: (2025)
Unsupervised Adaptive Deep Learning Method For BCI Motor Imagery Decoding
by: Ouahidi, Yassine El, et al.
Published: (2024)
by: Ouahidi, Yassine El, et al.
Published: (2024)
AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Towards a General Time Series Anomaly Detector with Adaptive Bottlenecks and Dual Adversarial Decoders
by: Shentu, Qichao, et al.
Published: (2024)
by: Shentu, Qichao, et al.
Published: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Computable Model-Independent Bounds for Adversarial Quantum Machine Learning
by: Li, Bacui, et al.
Published: (2024)
by: Li, Bacui, et al.
Published: (2024)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
by: Neelam, Sanjit, et al.
Published: (2025)
by: Neelam, Sanjit, et al.
Published: (2025)
Adaptive Resolving Methods for Reinforcement Learning with Function Approximations
by: Jiang, Jiashuo, et al.
Published: (2025)
by: Jiang, Jiashuo, et al.
Published: (2025)
TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
by: Li, Ye, et al.
Published: (2025)
by: Li, Ye, et al.
Published: (2025)
Similar Items
-
Intent-aligned Formal Specification Synthesis via Traceable Refinement
by: Ye, Zhe, et al.
Published: (2026) -
Sample-Optimal Agnostic Boosting with Unlabeled Data
by: Ghai, Udaya, et al.
Published: (2025) -
Sample-Efficient Agnostic Boosting
by: Ghai, Udaya, et al.
Published: (2024) -
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026) -
Replicable Bandits with UCB based Exploration
by: Deb, Rohan, et al.
Published: (2026)