Optimization of Epsilon-Greedy Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Che, Ethan, Ceylan, Hakan, McInerney, James, Kallus, Nathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
von: McInerney, James, et al.
Veröffentlicht: (2024)
von: McInerney, James, et al.
Veröffentlicht: (2024)
Entropy After </Think> for reasoning model early exiting
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Adjusting Regression Models for Conditional Uncertainty Calibration
von: Gao, Ruijiang, et al.
Veröffentlicht: (2024)
von: Gao, Ruijiang, et al.
Veröffentlicht: (2024)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
Epsilon-Greedy Thompson Sampling to Bayesian Optimization
von: Do, Bach, et al.
Veröffentlicht: (2024)
von: Do, Bach, et al.
Veröffentlicht: (2024)
A Statistical-Modelling Approach to Feedforward Neural Network Model Selection
von: McInerney, Andrew, et al.
Veröffentlicht: (2022)
von: McInerney, Andrew, et al.
Veröffentlicht: (2022)
Exploration in the Limit
von: Cho, Brian M., et al.
Veröffentlicht: (2025)
von: Cho, Brian M., et al.
Veröffentlicht: (2025)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
von: Kallus, Nathan
Veröffentlicht: (2025)
von: Kallus, Nathan
Veröffentlicht: (2025)
Accelerating Matrix Diagonalization through Decision Transformers with Epsilon-Greedy Optimization
von: Bhatta, Kshitij, et al.
Veröffentlicht: (2024)
von: Bhatta, Kshitij, et al.
Veröffentlicht: (2024)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
von: Cho, Brian, et al.
Veröffentlicht: (2024)
von: Cho, Brian, et al.
Veröffentlicht: (2024)
Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data
von: Oprescu, Miruna, et al.
Veröffentlicht: (2024)
von: Oprescu, Miruna, et al.
Veröffentlicht: (2024)
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
von: Kallus, Nathan, et al.
Veröffentlicht: (2020)
von: Kallus, Nathan, et al.
Veröffentlicht: (2020)
Robust and Agnostic Learning of Conditional Distributional Treatment Effects
von: Kallus, Nathan, et al.
Veröffentlicht: (2022)
von: Kallus, Nathan, et al.
Veröffentlicht: (2022)
Contextual Linear Optimization with Partial Feedback
von: Hu, Yichun, et al.
Veröffentlicht: (2024)
von: Hu, Yichun, et al.
Veröffentlicht: (2024)
A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
Demistifying Inference after Adaptive Experiments
von: Bibaut, Aurélien, et al.
Veröffentlicht: (2024)
von: Bibaut, Aurélien, et al.
Veröffentlicht: (2024)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Multi-Armed Bandits with Interference
von: Jia, Su, et al.
Veröffentlicht: (2024)
von: Jia, Su, et al.
Veröffentlicht: (2024)
Bellman Calibration for $V$-Learning in Offline Reinforcement Learning
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams
von: Cho, Brian, et al.
Veröffentlicht: (2024)
von: Cho, Brian, et al.
Veröffentlicht: (2024)
Teaching Knowledge Management (SIG KM).
von: McInerney, Claire
Veröffentlicht: (2000)
von: McInerney, Claire
Veröffentlicht: (2000)
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
von: Li, Tong, et al.
Veröffentlicht: (2026)
von: Li, Tong, et al.
Veröffentlicht: (2026)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2026)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2026)
Low-Rank MDPs with Continuous Action Spaces
von: Bennett, Andrew, et al.
Veröffentlicht: (2023)
von: Bennett, Andrew, et al.
Veröffentlicht: (2023)
The Central Role of the Loss Function in Reinforcement Learning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
Is Cosine-Similarity of Embeddings Really About Similarity?
von: Steck, Harald, et al.
Veröffentlicht: (2024)
von: Steck, Harald, et al.
Veröffentlicht: (2024)
Efficient Adaptive Experimentation with Noncompliance
von: Oprescu, Miruna, et al.
Veröffentlicht: (2025)
von: Oprescu, Miruna, et al.
Veröffentlicht: (2025)
Simulation-Based Inference for Adaptive Experiments
von: Cho, Brian M, et al.
Veröffentlicht: (2025)
von: Cho, Brian M, et al.
Veröffentlicht: (2025)
GAAVI: Global Asymptotic Anytime Valid Inference for the Conditional Mean Function
von: Cho, Brian M, et al.
Veröffentlicht: (2026)
von: Cho, Brian M, et al.
Veröffentlicht: (2026)
Functional Natural Policy Gradients
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2026)
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2026)
Clustered Switchback Designs for Experimentation Under Spatio-temporal Interference
von: Jia, Su, et al.
Veröffentlicht: (2023)
von: Jia, Su, et al.
Veröffentlicht: (2023)
Near-Optimal Non-Parametric Sequential Tests and Confidence Sequences with Possibly Dependent Observations
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2022)
von: Bibaut, Aurelien, et al.
Veröffentlicht: (2022)
Environmental Scanning and the Information Manager.
von: Newsome, James, et al.
Veröffentlicht: (1990)
von: Newsome, James, et al.
Veröffentlicht: (1990)
Greedy Alignment Principle for Optimizer Selection
von: Lee, Jaerin, et al.
Veröffentlicht: (2025)
von: Lee, Jaerin, et al.
Veröffentlicht: (2025)
Causal Inference on Networks under Misspecified Exposure Mappings: A Partial Identification Framework
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Long-term Causal Inference Under Persistent Confounding via Data Combination
von: Imbens, Guido, et al.
Veröffentlicht: (2022)
von: Imbens, Guido, et al.
Veröffentlicht: (2022)
Optimization-Driven Adaptive Experimentation
von: Che, Ethan, et al.
Veröffentlicht: (2024)
von: Che, Ethan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
von: McInerney, James, et al.
Veröffentlicht: (2024) -
Entropy After </Think> for reasoning model early exiting
von: Wang, Xi, et al.
Veröffentlicht: (2025) -
Adjusting Regression Models for Conditional Uncertainty Calibration
von: Gao, Ruijiang, et al.
Veröffentlicht: (2024) -
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024) -
Epsilon-Greedy Thompson Sampling to Bayesian Optimization
von: Do, Bach, et al.
Veröffentlicht: (2024)