Saved in:
| Main Authors: | Barbier-Chebbah, Alex, Vestergaard, Christian L., Masson, Jean-Baptiste, Boursier, Etienne |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2310.12563 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Information maximization for a broad variety of multi-armed bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2025)
by: Barbier-Chebbah, Alex, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Penalising the biases in norm regularisation enforces sparsity
by: Boursier, Etienne, et al.
Published: (2023)
by: Boursier, Etienne, et al.
Published: (2023)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Early alignment in two-layer networks training is a two-edged sword
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Revealing graph bandits for maximizing local influence
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Leveraging heterogeneous spillover in maximizing contextual bandit rewards
by: Faruk, Ahmed Sayeed, et al.
Published: (2023)
by: Faruk, Ahmed Sayeed, et al.
Published: (2023)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Benignity of loss landscape with weight decay requires both large overparametrization and initialization
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Truthful mechanisms for linear bandit games with private contexts
by: Hu, Yiting, et al.
Published: (2025)
by: Hu, Yiting, et al.
Published: (2025)
Fairness in two-player zero-sum games with bandit feedback
by: Akash, S, et al.
Published: (2026)
by: Akash, S, et al.
Published: (2026)
Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
by: Town, James, et al.
Published: (2026)
by: Town, James, et al.
Published: (2026)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Compression-based inference of network motif sets
by: Bénichou, Alexis, et al.
Published: (2023)
by: Bénichou, Alexis, et al.
Published: (2023)
Extreme bandits
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Optimal transport unlocks end-to-end learning for single-molecule localization
by: Seailles, Romain, et al.
Published: (2025)
by: Seailles, Romain, et al.
Published: (2025)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more
by: Kikkawa, Nobuaki, et al.
Published: (2024)
by: Kikkawa, Nobuaki, et al.
Published: (2024)
Online Decision-Focused Learning
by: Capitaine, Aymeric, et al.
Published: (2025)
by: Capitaine, Aymeric, et al.
Published: (2025)
Active clustering with bandit feedback
by: Thuot, Victor, et al.
Published: (2024)
by: Thuot, Victor, et al.
Published: (2024)
Minimum mean-squared error estimation with bandit feedback
by: Ghosh, Ayon, et al.
Published: (2022)
by: Ghosh, Ayon, et al.
Published: (2022)
Online learning in bandits with predicted context
by: Guo, Yongyi, et al.
Published: (2023)
by: Guo, Yongyi, et al.
Published: (2023)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Optimal Design for Reward Modeling in RLHF
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Risk and optimal policies in bandit experiments
by: Adusumilli, Karun
Published: (2021)
by: Adusumilli, Karun
Published: (2021)
Low-rank adaptive physics-informed HyperDeepONets for solving differential equations
by: Zeudong, Etienne, et al.
Published: (2025)
by: Zeudong, Etienne, et al.
Published: (2025)
On the optimal regret of collaborative personalized linear bandits
by: Huang, Bruce, et al.
Published: (2025)
by: Huang, Bruce, et al.
Published: (2025)
Offline-to-online hyperparameter transfer for stochastic bandits
by: Sharma, Dravyansh, et al.
Published: (2025)
by: Sharma, Dravyansh, et al.
Published: (2025)
Learning to Mitigate Externalities: the Coase Theorem with Hindsight Rationality
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Linear bandits with polylogarithmic minimax regret
by: Lumbreras, Josep, et al.
Published: (2024)
by: Lumbreras, Josep, et al.
Published: (2024)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
VITS : Variational Inference Thompson Sampling for contextual bandits
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Lookahead identification in adversarial bandits: accuracy and memory bounds
by: Brukhim, Nataly, et al.
Published: (2026)
by: Brukhim, Nataly, et al.
Published: (2026)
Efficient kernelized bandit algorithms via exploration distributions
by: Hu, Bingshan, et al.
Published: (2025)
by: Hu, Bingshan, et al.
Published: (2025)
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
Similar Items
-
Information maximization for a broad variety of multi-armed bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2025) -
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022) -
Penalising the biases in norm regularisation enforces sparsity
by: Boursier, Etienne, et al.
Published: (2023) -
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025) -
Early alignment in two-layer networks training is a two-edged sword
by: Boursier, Etienne, et al.
Published: (2024)