Learning Adaptive LLM Decoding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Su, Chloe H., Ye, Zhe, Tenka, Samuel, Yang, Aidan, Kong, Soonho, Ghai, Udaya |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Intent-aligned Formal Specification Synthesis via Traceable Refinement
par: Ye, Zhe, et autres
Publié: (2026)
par: Ye, Zhe, et autres
Publié: (2026)
Sample-Optimal Agnostic Boosting with Unlabeled Data
par: Ghai, Udaya, et autres
Publié: (2025)
par: Ghai, Udaya, et autres
Publié: (2025)
Sample-Efficient Agnostic Boosting
par: Ghai, Udaya, et autres
Publié: (2024)
par: Ghai, Udaya, et autres
Publié: (2024)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
par: Flamant, Cedric, et autres
Publié: (2026)
par: Flamant, Cedric, et autres
Publié: (2026)
Replicable Bandits with UCB based Exploration
par: Deb, Rohan, et autres
Publié: (2026)
par: Deb, Rohan, et autres
Publié: (2026)
Outbound Modeling for Inventory Management
par: Savorgnan, Riccardo, et autres
Publié: (2025)
par: Savorgnan, Riccardo, et autres
Publié: (2025)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
par: Song, Yuda, et autres
Publié: (2024)
par: Song, Yuda, et autres
Publié: (2024)
Neural Coordination and Capacity Control for Inventory Management
par: Eisenach, Carson, et autres
Publié: (2024)
par: Eisenach, Carson, et autres
Publié: (2024)
Adaptive Policy Learning Under Unknown Network Interference
par: Gleich, Aidan, et autres
Publié: (2026)
par: Gleich, Aidan, et autres
Publié: (2026)
Residual Learning and Context Encoding for Adaptive Offline-to-Online Reinforcement Learning
par: Nakhaei, Mohammadreza, et autres
Publié: (2024)
par: Nakhaei, Mohammadreza, et autres
Publié: (2024)
How Does Critical Batch Size Scale in Pre-training?
par: Zhang, Hanlin, et autres
Publié: (2024)
par: Zhang, Hanlin, et autres
Publié: (2024)
Teaching LLMs Program Semantics via Symbolic Execution Traces
par: Bayer, Jonas, et autres
Publié: (2026)
par: Bayer, Jonas, et autres
Publié: (2026)
Interactive Diabetes Risk Prediction Using Explainable Machine Learning: A Dash-Based Approach with SHAP, LIME, and Comorbidity Insights
par: Allani, Udaya
Publié: (2025)
par: Allani, Udaya
Publié: (2025)
MADQRL: Distributed Quantum Reinforcement Learning Framework for Multi-Agent Environments
par: Sawaika, Abhishek, et autres
Publié: (2026)
par: Sawaika, Abhishek, et autres
Publié: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
par: Kong, Linghao, et autres
Publié: (2026)
par: Kong, Linghao, et autres
Publié: (2026)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
par: Gao, Lei, et autres
Publié: (2025)
par: Gao, Lei, et autres
Publié: (2025)
Few-Shot Concept Unlearning with Low Rank Adaptation
par: Shreyas, Udaya, et autres
Publié: (2025)
par: Shreyas, Udaya, et autres
Publié: (2025)
Adaptive Budget Allocation in LLM-Augmented Surveys
par: Ye, Zikun, et autres
Publié: (2026)
par: Ye, Zikun, et autres
Publié: (2026)
Stroke Prediction using Clinical and Social Features in Machine Learning
par: Chadha, Aidan
Publié: (2024)
par: Chadha, Aidan
Publié: (2024)
Learning from Similar Linear Representations: Adaptivity, Minimaxity, and Robustness
par: Tian, Ye, et autres
Publié: (2023)
par: Tian, Ye, et autres
Publié: (2023)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
par: Lu, Jialin, et autres
Publié: (2026)
par: Lu, Jialin, et autres
Publié: (2026)
Tokenized Bandit for LLM Decoding and Alignment
par: Shin, Suho, et autres
Publié: (2025)
par: Shin, Suho, et autres
Publié: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
par: Wang, Meixuan, et autres
Publié: (2025)
par: Wang, Meixuan, et autres
Publié: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
par: Xu, WeiZhe, et autres
Publié: (2026)
par: Xu, WeiZhe, et autres
Publié: (2026)
Entropy Regularized Task Representation Learning for Offline Meta-Reinforcement Learning
par: Nakhaei, Mohammadreza, et autres
Publié: (2024)
par: Nakhaei, Mohammadreza, et autres
Publié: (2024)
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
par: Ye, Zicong, et autres
Publié: (2025)
par: Ye, Zicong, et autres
Publié: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
par: Yang, Seongjun, et autres
Publié: (2023)
par: Yang, Seongjun, et autres
Publié: (2023)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
par: Jaiswal, Ajay, et autres
Publié: (2024)
par: Jaiswal, Ajay, et autres
Publié: (2024)
ElfCore: A 28nm Neural Processor Enabling Dynamic Structured Sparse Training and Online Self-Supervised Learning with Activity-Dependent Weight Update
par: Su, Zhe, et autres
Publié: (2025)
par: Su, Zhe, et autres
Publié: (2025)
BRIDGE: Building Representations In Domain Guided Program Synthesis
par: George, Robert Joseph, et autres
Publié: (2025)
par: George, Robert Joseph, et autres
Publié: (2025)
Unsupervised Adaptive Deep Learning Method For BCI Motor Imagery Decoding
par: Ouahidi, Yassine El, et autres
Publié: (2024)
par: Ouahidi, Yassine El, et autres
Publié: (2024)
AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
par: Zhou, Hongyi, et autres
Publié: (2025)
par: Zhou, Hongyi, et autres
Publié: (2025)
Towards a General Time Series Anomaly Detector with Adaptive Bottlenecks and Dual Adversarial Decoders
par: Shentu, Qichao, et autres
Publié: (2024)
par: Shentu, Qichao, et autres
Publié: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
par: Qiao, Ye, et autres
Publié: (2025)
par: Qiao, Ye, et autres
Publié: (2025)
Computable Model-Independent Bounds for Adversarial Quantum Machine Learning
par: Li, Bacui, et autres
Publié: (2024)
par: Li, Bacui, et autres
Publié: (2024)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
par: Ramakrishnan, Ramchalam Kinattinkara, et autres
Publié: (2025)
par: Ramakrishnan, Ramchalam Kinattinkara, et autres
Publié: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
par: Yang, Lijie, et autres
Publié: (2024)
par: Yang, Lijie, et autres
Publié: (2024)
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
par: Neelam, Sanjit, et autres
Publié: (2025)
par: Neelam, Sanjit, et autres
Publié: (2025)
Adaptive Resolving Methods for Reinforcement Learning with Function Approximations
par: Jiang, Jiashuo, et autres
Publié: (2025)
par: Jiang, Jiashuo, et autres
Publié: (2025)
TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
par: Li, Ye, et autres
Publié: (2025)
par: Li, Ye, et autres
Publié: (2025)
Documents similaires
-
Intent-aligned Formal Specification Synthesis via Traceable Refinement
par: Ye, Zhe, et autres
Publié: (2026) -
Sample-Optimal Agnostic Boosting with Unlabeled Data
par: Ghai, Udaya, et autres
Publié: (2025) -
Sample-Efficient Agnostic Boosting
par: Ghai, Udaya, et autres
Publié: (2024) -
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
par: Flamant, Cedric, et autres
Publié: (2026) -
Replicable Bandits with UCB based Exploration
par: Deb, Rohan, et autres
Publié: (2026)