Finite-Time Logarithmic Bayes Regret Upper Bounds
Fuente:
arXiv
Salvato in:
| Autori principali: | Atsidakou, Alexia, Kveton, Branislav, Katariya, Sumeet, Caramanis, Constantine, Sanghavi, Sujay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Asymptotically-Optimal Gaussian Bandits with Side Observations
di: Atsidakou, Alexia, et al.
Pubblicazione: (2025)
di: Atsidakou, Alexia, et al.
Pubblicazione: (2025)
Contextual Pandora's Box
di: Atsidakou, Alexia, et al.
Pubblicazione: (2022)
di: Atsidakou, Alexia, et al.
Pubblicazione: (2022)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
Selective Uncertainty Propagation in Offline RL
di: Krishnamurthy, Sanath Kumar, et al.
Pubblicazione: (2023)
di: Krishnamurthy, Sanath Kumar, et al.
Pubblicazione: (2023)
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Context-Free Synthetic Data Mitigates Forgetting
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
Optimization Can Learn Johnson Lindenstrauss Embeddings
di: Tsikouras, Nikos, et al.
Pubblicazione: (2024)
di: Tsikouras, Nikos, et al.
Pubblicazione: (2024)
On Mitigating Affinity Bias through Bandits with Evolving Biased Feedback
di: Faw, Matthew, et al.
Pubblicazione: (2025)
di: Faw, Matthew, et al.
Pubblicazione: (2025)
MaxSketch: Robust Distinct Counting in Streams via Random Projections
di: Tsikouras, Nikos, et al.
Pubblicazione: (2026)
di: Tsikouras, Nikos, et al.
Pubblicazione: (2026)
Test-Time Speculation
di: Kumar, Avinash, et al.
Pubblicazione: (2026)
di: Kumar, Avinash, et al.
Pubblicazione: (2026)
Enabling Approximate Joint Sampling in Diffusion LMs
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
Anchored Diffusion Language Model
di: Rout, Litu, et al.
Pubblicazione: (2025)
di: Rout, Litu, et al.
Pubblicazione: (2025)
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
LLM-as-Judge on a Budget
di: Saha, Aadirupa, et al.
Pubblicazione: (2026)
di: Saha, Aadirupa, et al.
Pubblicazione: (2026)
Cross-Validated Off-Policy Evaluation
di: Cief, Matej, et al.
Pubblicazione: (2024)
di: Cief, Matej, et al.
Pubblicazione: (2024)
Learning Mixtures of Experts with EM: A Mirror Descent Perspective
di: Fruytier, Quentin, et al.
Pubblicazione: (2024)
di: Fruytier, Quentin, et al.
Pubblicazione: (2024)
Understanding Self-Supervised Learning via Gaussian Mixture Models
di: Bansal, Parikshit, et al.
Pubblicazione: (2024)
di: Bansal, Parikshit, et al.
Pubblicazione: (2024)
Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
AnCoder: Anchored Code Generation via Discrete Diffusion Models
di: Xue, Anton, et al.
Pubblicazione: (2026)
di: Xue, Anton, et al.
Pubblicazione: (2026)
Geometric Median (GM) Matching for Robust Data Pruning
di: Acharya, Anish, et al.
Pubblicazione: (2024)
di: Acharya, Anish, et al.
Pubblicazione: (2024)
Logarithmic Regret for Nonlinear Control
di: Wang, James, et al.
Pubblicazione: (2025)
di: Wang, James, et al.
Pubblicazione: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
di: Cief, Matej, et al.
Pubblicazione: (2022)
di: Cief, Matej, et al.
Pubblicazione: (2022)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
di: Kwon, Jeongyeol, et al.
Pubblicazione: (2024)
di: Kwon, Jeongyeol, et al.
Pubblicazione: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
Semi-supervised learning with max-margin graph cuts
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Spectral bandits for smooth graph functions
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Online semi-supervised perception: Real-time learning without explicit feedback
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Time Weaver: A Conditional Time Series Generation Model
di: Narasimhan, Sai Shankar, et al.
Pubblicazione: (2024)
di: Narasimhan, Sai Shankar, et al.
Pubblicazione: (2024)
Linear Regression with Unknown Truncation Beyond Gaussian Features
di: Kouridakis, Alexandros, et al.
Pubblicazione: (2026)
di: Kouridakis, Alexandros, et al.
Pubblicazione: (2026)
Blocking Bandits
di: Basu, Soumya, et al.
Pubblicazione: (2019)
di: Basu, Soumya, et al.
Pubblicazione: (2019)
Improved Regret Bounds for Gaussian Process Upper Confidence Bound in Bayesian Optimization
di: Iwazaki, Shogo
Pubblicazione: (2025)
di: Iwazaki, Shogo
Pubblicazione: (2025)
Regret Analysis for Randomized Gaussian Process Upper Confidence Bound
di: Takeno, Shion, et al.
Pubblicazione: (2024)
di: Takeno, Shion, et al.
Pubblicazione: (2024)
Online Inverse Linear Optimization: Efficient Logarithmic-Regret Algorithm, Robustness to Suboptimality, and Lower Bound
di: Sakaue, Shinsaku, et al.
Pubblicazione: (2025)
di: Sakaue, Shinsaku, et al.
Pubblicazione: (2025)
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
di: Acharya, Anish, et al.
Pubblicazione: (2025)
di: Acharya, Anish, et al.
Pubblicazione: (2025)
Regretful Decisions under Label Noise
di: Nagaraj, Sujay, et al.
Pubblicazione: (2025)
di: Nagaraj, Sujay, et al.
Pubblicazione: (2025)
Spectral bandits for smooth graph functions with applications in recommender systems
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
Off-Policy Evaluation from Logged Human Feedback
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
Online Posterior Sampling with a Diffusion Prior
di: Kveton, Branislav, et al.
Pubblicazione: (2024)
di: Kveton, Branislav, et al.
Pubblicazione: (2024)
Evidence-based anomaly detection in clinical domains
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Asymptotically-Optimal Gaussian Bandits with Side Observations
di: Atsidakou, Alexia, et al.
Pubblicazione: (2025) -
Contextual Pandora's Box
di: Atsidakou, Alexia, et al.
Pubblicazione: (2022) -
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026) -
Selective Uncertainty Propagation in Offline RL
di: Krishnamurthy, Sanath Kumar, et al.
Pubblicazione: (2023) -
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026)