Feel-Good Thompson Sampling for Contextual Bandits: a Markov Chain Monte Carlo Showdown
Fuente:
arXiv
Saved in:
| Main Authors: | Anand, Emile, Liaw, Sarah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
by: Zhang, Zhipeng, et al.
Published: (2026)
by: Zhang, Zhipeng, et al.
Published: (2026)
Implicit Counterfactual Data Augmentation for Robust Learning
by: Zhou, Xiaoling, et al.
Published: (2023)
by: Zhou, Xiaoling, et al.
Published: (2023)
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
by: Leyli-abadi, Milad, et al.
Published: (2025)
by: Leyli-abadi, Milad, et al.
Published: (2025)
Graceful task adaptation with a bi-hemispheric RL agent
by: Nicholas, Grant, et al.
Published: (2024)
by: Nicholas, Grant, et al.
Published: (2024)
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
by: Elliker, Clément, et al.
Published: (2025)
by: Elliker, Clément, et al.
Published: (2025)
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
by: Fouladi, Kasra, et al.
Published: (2026)
by: Fouladi, Kasra, et al.
Published: (2026)
An Axiomatic Approach to General Intelligence: SANC(E3) -- Self-organizing Active Network of Concepts with Energy E3
by: Kwon, Daesuk, et al.
Published: (2026)
by: Kwon, Daesuk, et al.
Published: (2026)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Uncertainty Quantification Using Ensemble Learning and Monte Carlo Sampling for Performance Prediction and Monitoring in Cell Culture Processes
by: Khuat, Thanh Tung, et al.
Published: (2024)
by: Khuat, Thanh Tung, et al.
Published: (2024)
Choosing DAG Models Using Markov and Minimal Edge Count in the Absence of Ground Truth
by: Ramsey, Joseph D., et al.
Published: (2024)
by: Ramsey, Joseph D., et al.
Published: (2024)
CoGraM: Context-sensitive granular optimization method with rollback for robust model fusion
by: Lenz, Julius
Published: (2025)
by: Lenz, Julius
Published: (2025)
Adaptive Latent-Space Constraints in Personalized Federated Learning
by: Ayromlou, Sana, et al.
Published: (2025)
by: Ayromlou, Sana, et al.
Published: (2025)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
by: Vahdati, Sahar, et al.
Published: (2026)
by: Vahdati, Sahar, et al.
Published: (2026)
Deep Reinforcement Learning for Adverse Garage Scenario Generation
by: Li, Kai
Published: (2024)
by: Li, Kai
Published: (2024)
AutoHood3D: A Multi-Modal Benchmark for Automotive Hood Design and Fluid-Structure Interaction
by: Sharma, Vansh, et al.
Published: (2025)
by: Sharma, Vansh, et al.
Published: (2025)
The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
by: Lovén, Lauri, et al.
Published: (2026)
by: Lovén, Lauri, et al.
Published: (2026)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
LaPro-DTA: Latent Dual-View Drug Representations and Salient Protein Feature Extraction for Generalizable Drug--Target Affinity Prediction
by: Dun, Zihan, et al.
Published: (2026)
by: Dun, Zihan, et al.
Published: (2026)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023)
by: Ferreira, Alexandre R., et al.
Published: (2023)
Intervention Complexity as a Canonical Reward and a Measure of Intelligence
by: McCane, Brendan
Published: (2026)
by: McCane, Brendan
Published: (2026)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
by: Vegner, Ivan, et al.
Published: (2025)
by: Vegner, Ivan, et al.
Published: (2025)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
by: Baxi, Rahul
Published: (2025)
by: Baxi, Rahul
Published: (2025)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
by: Kaplanski, Pawel
Published: (2026)
by: Kaplanski, Pawel
Published: (2026)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
by: Minai, Ali A.
Published: (2025)
by: Minai, Ali A.
Published: (2025)
Beyond Mimicry: Preference Coherence in LLMs
by: Mikaelson, Luhan, et al.
Published: (2025)
by: Mikaelson, Luhan, et al.
Published: (2025)
Prompt Readiness Levels (PRL): a maturity scale and scoring framework for production grade prompt assets
by: Guinard, Sebastien
Published: (2026)
by: Guinard, Sebastien
Published: (2026)
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
by: Ma, Jingxiao, et al.
Published: (2025)
by: Ma, Jingxiao, et al.
Published: (2025)
The Blueprints of Intelligence: A Functional-Topological Foundation for Perception and Representation
by: Di Santi, Eduardo
Published: (2025)
by: Di Santi, Eduardo
Published: (2025)
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
by: Wang, Xiaoxiao, et al.
Published: (2026)
by: Wang, Xiaoxiao, et al.
Published: (2026)
A Survey on Data-Dependent Worst-Case Generalization Bounds
by: Leroux, Hubert, et al.
Published: (2026)
by: Leroux, Hubert, et al.
Published: (2026)
5G Traffic Prediction with Time Series Analysis
by: Nayak, Nikhil, et al.
Published: (2021)
by: Nayak, Nikhil, et al.
Published: (2021)
Uniform $\mathcal{C}^k$ Approximation of $G$-Invariant and Antisymmetric Functions, Embedding Dimensions, and Polynomial Representations
by: Ganguly, Soumya, et al.
Published: (2024)
by: Ganguly, Soumya, et al.
Published: (2024)
Error-related Potential driven Reinforcement Learning for adaptive Brain-Computer Interfaces
by: Fidêncio, Aline Xavier, et al.
Published: (2025)
by: Fidêncio, Aline Xavier, et al.
Published: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
by: Hooshyar, Danial, et al.
Published: (2023)
by: Hooshyar, Danial, et al.
Published: (2023)
Improving Fairness with Ensemble Combination: Margin-Dependent Bounds
by: Bian, Yijun
Published: (2023)
by: Bian, Yijun
Published: (2023)
What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs
by: Saqur, Raeid
Published: (2024)
by: Saqur, Raeid
Published: (2024)
Similar Items
-
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
by: Zhang, Zhipeng, et al.
Published: (2026) -
Implicit Counterfactual Data Augmentation for Robust Learning
by: Zhou, Xiaoling, et al.
Published: (2023) -
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
by: Leyli-abadi, Milad, et al.
Published: (2025) -
Graceful task adaptation with a bi-hemispheric RL agent
by: Nicholas, Grant, et al.
Published: (2024) -
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
by: Elliker, Clément, et al.
Published: (2025)