Tokenized Bandit for LLM Decoding and Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Suho, Yang, Chenghao, Xu, Haifeng, Hajiaghayi, Mohammad T. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ad Auctions for LLMs via Retrieval Augmented Generation
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
Replication-proof Bandit Mechanism Design with Bayesian Agents
by: Shin, Suho, et al.
Published: (2023)
by: Shin, Suho, et al.
Published: (2023)
Bandit Social Learning: Exploration under Myopic Behavior
by: Banihashem, Kiarash, et al.
Published: (2023)
by: Banihashem, Kiarash, et al.
Published: (2023)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
Online Advertisements with LLMs: Opportunities and Challenges
by: Feizi, Soheil, et al.
Published: (2023)
by: Feizi, Soheil, et al.
Published: (2023)
Regret Analysis of Repeated Delegated Choice
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2023)
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2023)
Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding
by: Jing, Zihao, et al.
Published: (2026)
by: Jing, Zihao, et al.
Published: (2026)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Online Domain-aware LLM Decoding for Continual Domain Evolution
by: Abu-Shaira, Mohammad, et al.
Published: (2026)
by: Abu-Shaira, Mohammad, et al.
Published: (2026)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
by: Zou, Mingxi, et al.
Published: (2026)
by: Zou, Mingxi, et al.
Published: (2026)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024)
by: Deb, Rohan, et al.
Published: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
by: Pedramfar, Mohammad, et al.
Published: (2023)
by: Pedramfar, Mohammad, et al.
Published: (2023)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
by: Pershin, Maksim, et al.
Published: (2026)
by: Pershin, Maksim, et al.
Published: (2026)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Dueling Over Dessert, Mastering the Art of Repeated Cake Cutting
by: Brânzei, Simina, et al.
Published: (2024)
by: Brânzei, Simina, et al.
Published: (2024)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
by: Wang, Duo, et al.
Published: (2024)
by: Wang, Duo, et al.
Published: (2024)
Token-Efficient RL for LLM Reasoning
by: Lee, Alan, et al.
Published: (2025)
by: Lee, Alan, et al.
Published: (2025)
Adversarial Preference Learning for Robust LLM Alignment
by: Wang, Yuanfu, et al.
Published: (2025)
by: Wang, Yuanfu, et al.
Published: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023)
by: Liu, Xutong, et al.
Published: (2023)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
Tracing Mathematical Proficiency Through Problem-Solving Processes
by: Park, Jungyang, et al.
Published: (2025)
by: Park, Jungyang, et al.
Published: (2025)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
by: Kim, Jung-hun, et al.
Published: (2017)
by: Kim, Jung-hun, et al.
Published: (2017)
Algorithmic Delegated Choice: An Annotated Reading List
by: Hajiaghayi, Mohammad T., et al.
Published: (2025)
by: Hajiaghayi, Mohammad T., et al.
Published: (2025)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
by: Ran, Dezhi, et al.
Published: (2025)
by: Ran, Dezhi, et al.
Published: (2025)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026)
by: Li, Yingru, et al.
Published: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
by: Wang, Jiaxuan, et al.
Published: (2026)
by: Wang, Jiaxuan, et al.
Published: (2026)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
by: Khosravi, Hamed, et al.
Published: (2025)
by: Khosravi, Hamed, et al.
Published: (2025)
EPLKG: Efficient Prompt Learning with Knowledge Graph
by: Lim, YongTaek, et al.
Published: (2023)
by: Lim, YongTaek, et al.
Published: (2023)
Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits
by: Bayley, Adam, et al.
Published: (2026)
by: Bayley, Adam, et al.
Published: (2026)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
by: Kim, Taesan, et al.
Published: (2026)
by: Kim, Taesan, et al.
Published: (2026)
Similar Items
-
Ad Auctions for LLMs via Retrieval Augmented Generation
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024) -
Replication-proof Bandit Mechanism Design with Bayesian Agents
by: Shin, Suho, et al.
Published: (2023) -
Bandit Social Learning: Exploration under Myopic Behavior
by: Banihashem, Kiarash, et al.
Published: (2023) -
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
by: Yang, Chenghao, et al.
Published: (2025) -
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)