Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yingru, Luo, Zhi-Quan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Exploration via Ensemble++
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
Fixed-Budget Differentially Private Best Arm Identification
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
Universal time-series forecasting with mixture predictors
von: Ryabko, Daniil
Veröffentlicht: (2020)
von: Ryabko, Daniil
Veröffentlicht: (2020)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
von: Tohme, Tony, et al.
Veröffentlicht: (2023)
von: Tohme, Tony, et al.
Veröffentlicht: (2023)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
Entropy, concentration, and learning: a statistical mechanics primer
von: Balsubramani, Akshay
Veröffentlicht: (2024)
von: Balsubramani, Akshay
Veröffentlicht: (2024)
Probability Tools for Sequential Random Projection
von: Li, Yingru
Veröffentlicht: (2024)
von: Li, Yingru
Veröffentlicht: (2024)
High-probability sample complexities for policy evaluation with linear function approximation
von: Li, Gen, et al.
Veröffentlicht: (2023)
von: Li, Gen, et al.
Veröffentlicht: (2023)
Influence functions and regularity tangents for efficient active learning
von: Eaton, Frederik
Veröffentlicht: (2024)
von: Eaton, Frederik
Veröffentlicht: (2024)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
Random Multiplexing
von: Liu, Lei, et al.
Veröffentlicht: (2025)
von: Liu, Lei, et al.
Veröffentlicht: (2025)
On the best approximation by finite Gaussian mixtures
von: Ma, Yun, et al.
Veröffentlicht: (2024)
von: Ma, Yun, et al.
Veröffentlicht: (2024)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
von: Levy, Jordan, et al.
Veröffentlicht: (2026)
von: Levy, Jordan, et al.
Veröffentlicht: (2026)
On the Separability of Information in Diffusion Models
von: Premkumar, Akhil
Veröffentlicht: (2025)
von: Premkumar, Akhil
Veröffentlicht: (2025)
Guaranteed Recovery of Unambiguous Clusters
von: Mazooji, Kayvon, et al.
Veröffentlicht: (2025)
von: Mazooji, Kayvon, et al.
Veröffentlicht: (2025)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
Statistical inference with belief functions: A survey
von: Cuzzolin, Fabio
Veröffentlicht: (2026)
von: Cuzzolin, Fabio
Veröffentlicht: (2026)
Generalising realisability in statistical learning theory under epistemic uncertainty
von: Cuzzolin, Fabio
Veröffentlicht: (2024)
von: Cuzzolin, Fabio
Veröffentlicht: (2024)
Cost-optimal Sequential Testing via Doubly Robust Q-learning
von: Zhou, Doudou, et al.
Veröffentlicht: (2026)
von: Zhou, Doudou, et al.
Veröffentlicht: (2026)
A comparative study of conformal prediction methods for valid uncertainty quantification in machine learning
von: Dewolf, Nicolas
Veröffentlicht: (2024)
von: Dewolf, Nicolas
Veröffentlicht: (2024)
Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
von: Mou, Wenlong
Veröffentlicht: (2026)
von: Mou, Wenlong
Veröffentlicht: (2026)
A non-asymptotic distributional theory of approximate message passing for sparse and robust regression
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Discovering highly efficient low-weight quantum error-correcting codes with reinforcement learning
von: He, Austin Yubo, et al.
Veröffentlicht: (2025)
von: He, Austin Yubo, et al.
Veröffentlicht: (2025)
Transport f divergences
von: Li, Wuchen
Veröffentlicht: (2025)
von: Li, Wuchen
Veröffentlicht: (2025)
Ridge interpolators in correlated factor regression models -- exact risk analysis
von: Stojnic, Mihailo
Veröffentlicht: (2024)
von: Stojnic, Mihailo
Veröffentlicht: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
von: Guo, Yang, et al.
Veröffentlicht: (2025)
von: Guo, Yang, et al.
Veröffentlicht: (2025)
Beyond identifiability: Learning causal representations with few environments and finite samples
von: Lee, Inbeom, et al.
Veröffentlicht: (2026)
von: Lee, Inbeom, et al.
Veröffentlicht: (2026)
Precise analysis of ridge interpolators under heavy correlations -- a Random Duality Theory view
von: Stojnic, Mihailo
Veröffentlicht: (2024)
von: Stojnic, Mihailo
Veröffentlicht: (2024)
Perturbative adaptive importance sampling for Bayesian LOO cross-validation
von: Chang, Joshua C, et al.
Veröffentlicht: (2024)
von: Chang, Joshua C, et al.
Veröffentlicht: (2024)
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Representation learning with a transformer by contrastive learning for money laundering detection
von: Guéneau, Harold, et al.
Veröffentlicht: (2025)
von: Guéneau, Harold, et al.
Veröffentlicht: (2025)
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
On the sample complexity of parameter estimation in logistic regression with normal design
von: Hsu, Daniel, et al.
Veröffentlicht: (2023)
von: Hsu, Daniel, et al.
Veröffentlicht: (2023)
Thompson sampling: Precise arm-pull dynamics and adaptive inference
von: Han, Qiyang
Veröffentlicht: (2026)
von: Han, Qiyang
Veröffentlicht: (2026)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scalable Exploration via Ensemble++
von: Li, Yingru, et al.
Veröffentlicht: (2024) -
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
von: Chen, Zhirui, et al.
Veröffentlicht: (2024) -
Fixed-Budget Differentially Private Best Arm Identification
von: Chen, Zhirui, et al.
Veröffentlicht: (2024) -
Universal time-series forecasting with mixture predictors
von: Ryabko, Daniil
Veröffentlicht: (2020) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)