Thompson Sampling via Fine-Tuning of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Menet, Nicolas, Terzić, Aleksandar, Hersche, Michael, Krause, Andreas, Rahimi, Abbas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2025)
by: Terzić, Aleksandar, et al.
Published: (2025)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
Locally Coherent Parallel Decoding in Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2026)
by: Hersche, Michael, et al.
Published: (2026)
Limits of Transformer Language Models on Learning to Compose Algorithms
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
by: Terzić, Aleksandar, et al.
Published: (2024)
by: Terzić, Aleksandar, et al.
Published: (2024)
Towards Learning Abductive Reasoning using VSA Distributed Representations
by: Camposampiero, Giacomo, et al.
Published: (2024)
by: Camposampiero, Giacomo, et al.
Published: (2024)
Terminating Differentiable Tree Experts
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
On the Role of Noise in Factorizers for Disentangling Distributed Representations
by: Karunaratne, Geethan, et al.
Published: (2024)
by: Karunaratne, Geethan, et al.
Published: (2024)
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
Probabilistic Abduction for Visual Abstract Reasoning via Learning Rules in Vector-symbolic Architectures
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
Scalable Evaluation and Neural Models for Compositional Generalization
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
A foundation model with multi-variate parallel attention to generate neuronal activity
by: Carzaniga, Francesco, et al.
Published: (2025)
by: Carzaniga, Francesco, et al.
Published: (2025)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
Soft-Masked Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2025)
by: Hersche, Michael, et al.
Published: (2025)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
The Case for Cleaner Biosignals: High-fidelity Neural Compressor Enables Transfer from Cleaner iEEG to Noisier EEG
by: Carzaniga, Francesco Stefano, et al.
Published: (2025)
by: Carzaniga, Francesco Stefano, et al.
Published: (2025)
Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
by: Namkoong, Hongseok, et al.
Published: (2020)
by: Namkoong, Hongseok, et al.
Published: (2020)
GIFT: Global stabilisation via Intrinsic Fine Tuning
by: Young, Rory, et al.
Published: (2026)
by: Young, Rory, et al.
Published: (2026)
Linearization Explains Fine-Tuning in Large Language Models
by: Afzal, Zahra Rahimi, et al.
Published: (2026)
by: Afzal, Zahra Rahimi, et al.
Published: (2026)
Contextual Thompson Sampling via Generation of Missing Data
by: Zhang, Kelly W., et al.
Published: (2025)
by: Zhang, Kelly W., et al.
Published: (2025)
Graph Neural Thompson Sampling
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
by: Liu, Zehua, et al.
Published: (2026)
by: Liu, Zehua, et al.
Published: (2026)
Understanding and Preserving Safety in Fine-Tuned LLMs
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
FedRTS: Federated Robust Pruning via Combinatorial Thompson Sampling
by: Huang, Hong, et al.
Published: (2025)
by: Huang, Hong, et al.
Published: (2025)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
by: Arteaga, Gabriel Y., et al.
Published: (2024)
by: Arteaga, Gabriel Y., et al.
Published: (2024)
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
by: Sanyal, Sunny, et al.
Published: (2025)
by: Sanyal, Sunny, et al.
Published: (2025)
Improving Thompson Sampling via Information Relaxation for Budgeted Multi-armed Bandits
by: Jeong, Woojin, et al.
Published: (2024)
by: Jeong, Woojin, et al.
Published: (2024)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
by: Yu, Ziming, et al.
Published: (2024)
by: Yu, Ziming, et al.
Published: (2024)
Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
by: Guo, Wentao, et al.
Published: (2024)
by: Guo, Wentao, et al.
Published: (2024)
Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
by: Lin, Guang, et al.
Published: (2026)
by: Lin, Guang, et al.
Published: (2026)
Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
MINTS: Minimalist Thompson Sampling
by: Wang, Kaizheng
Published: (2026)
by: Wang, Kaizheng
Published: (2026)
Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning
by: Tang, Zhenchao, et al.
Published: (2025)
by: Tang, Zhenchao, et al.
Published: (2025)
Federated Fine-Tuning of LLMs: Framework Comparison and Research Directions
by: Yan, Na, et al.
Published: (2025)
by: Yan, Na, et al.
Published: (2025)
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
by: Deng, Yanxia, et al.
Published: (2025)
by: Deng, Yanxia, et al.
Published: (2025)
DONOD: Efficient and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning
by: Hu, Jucheng, et al.
Published: (2025)
by: Hu, Jucheng, et al.
Published: (2025)
Similar Items
-
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2025) -
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026) -
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026) -
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
by: Menet, Nicolas, et al.
Published: (2026) -
Locally Coherent Parallel Decoding in Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2026)