Low-rank Matrix Bandits with Heavy-tailed Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Yue, Hsieh, Cho-Jui, Lee, Thomas C. M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems
von: Kang, Yue, et al.
Veröffentlicht: (2024)
von: Kang, Yue, et al.
Veröffentlicht: (2024)
Online Continuous Hyperparameter Optimization for Generalized Linear Contextual Bandits
von: Kang, Yue, et al.
Veröffentlicht: (2023)
von: Kang, Yue, et al.
Veröffentlicht: (2023)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Generalized Low-Rank Matrix Contextual Bandits with Graph Information
von: Wang, Yao, et al.
Veröffentlicht: (2025)
von: Wang, Yao, et al.
Veröffentlicht: (2025)
Lipschitz Bandits with Stochastic Delayed Feedback
von: Liu, Zhongxuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxuan, et al.
Veröffentlicht: (2025)
Multi-agent Multi-armed Bandit with Fully Heavy-tailed Dynamics
von: Wang, Xingyu, et al.
Veröffentlicht: (2025)
von: Wang, Xingyu, et al.
Veröffentlicht: (2025)
IRIS: Intrinsic Reward Image Synthesis
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
von: Li, Lichen, et al.
Veröffentlicht: (2026)
von: Li, Lichen, et al.
Veröffentlicht: (2026)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
von: Kang, Yue, et al.
Veröffentlicht: (2025)
von: Kang, Yue, et al.
Veröffentlicht: (2025)
Heavy-tailed Linear Bandits: Adversarial Robustness, Best-of-both-worlds, and Beyond
von: Zhao, Canzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Canzhe, et al.
Veröffentlicht: (2025)
Expert Proximity as Surrogate Rewards for Single Demonstration Imitation Learning
von: Chiang, Chia-Cheng, et al.
Veröffentlicht: (2024)
von: Chiang, Chia-Cheng, et al.
Veröffentlicht: (2024)
Adversarial Examples Detection with Bayesian Neural Network
von: Li, Yao, et al.
Veröffentlicht: (2021)
von: Li, Yao, et al.
Veröffentlicht: (2021)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
$t^3$-Variational Autoencoder: Learning Heavy-tailed Data with Student's t and Power Divergence
von: Kim, Juno, et al.
Veröffentlicht: (2023)
von: Kim, Juno, et al.
Veröffentlicht: (2023)
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2026)
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2026)
Data Attribution for Diffusion Models: Timestep-induced Bias in Influence Estimation
von: Xie, Tong, et al.
Veröffentlicht: (2024)
von: Xie, Tong, et al.
Veröffentlicht: (2024)
Quantum Lipschitz Bandits
von: Yi, Bongsoo, et al.
Veröffentlicht: (2025)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2025)
BanditQ: Fair Bandits with Guaranteed Rewards
von: Sinha, Abhishek
Veröffentlicht: (2023)
von: Sinha, Abhishek
Veröffentlicht: (2023)
Provably Robust Training of Quantum Circuit Classifiers Against Parameter Noise
von: Tecot, Lucas, et al.
Veröffentlicht: (2025)
von: Tecot, Lucas, et al.
Veröffentlicht: (2025)
Stochastic Low-rank Tensor Bandits for Multi-dimensional Online Decision Making
von: Zhou, Jie, et al.
Veröffentlicht: (2020)
von: Zhou, Jie, et al.
Veröffentlicht: (2020)
An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning
von: Bai, Andrew, et al.
Veröffentlicht: (2024)
von: Bai, Andrew, et al.
Veröffentlicht: (2024)
Generalizing Score-based generative models for Heavy-tailed Distributions
von: Fassina, Tiziano, et al.
Veröffentlicht: (2026)
von: Fassina, Tiziano, et al.
Veröffentlicht: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
von: Ye, Hao, et al.
Veröffentlicht: (2026)
von: Ye, Hao, et al.
Veröffentlicht: (2026)
On the Loss of Context-awareness in General Instruction Fine-tuning
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
Bandit Simulation for Average Reward Inference
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2024)
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2024)
Biased Dueling Bandits with Stochastic Delayed Feedback
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
FlexAct: Why Learn when you can Pick?
von: Kumar, Ramnath, et al.
Veröffentlicht: (2026)
von: Kumar, Ramnath, et al.
Veröffentlicht: (2026)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
von: Cho, Brian, et al.
Veröffentlicht: (2024)
von: Cho, Brian, et al.
Veröffentlicht: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
von: He, Jack, et al.
Veröffentlicht: (2024)
von: He, Jack, et al.
Veröffentlicht: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
Heavy-tailed Contamination is Easier than Adversarial Contamination
von: Cherapanamjeri, Yeshwanth, et al.
Veröffentlicht: (2024)
von: Cherapanamjeri, Yeshwanth, et al.
Veröffentlicht: (2024)
Online Minimization of Polarization and Disagreement via Low-Rank Matrix Bandits
von: Cinus, Federico, et al.
Veröffentlicht: (2025)
von: Cinus, Federico, et al.
Veröffentlicht: (2025)
Concept Gradient: Concept-based Interpretation Without Linear Assumption
von: Bai, Andrew, et al.
Veröffentlicht: (2022)
von: Bai, Andrew, et al.
Veröffentlicht: (2022)
Continual Learning of Numerous Tasks from Long-tail Distributions
von: Kang, Liwei, et al.
Veröffentlicht: (2024)
von: Kang, Liwei, et al.
Veröffentlicht: (2024)
Certified Training with Branch-and-Bound for Lyapunov-stable Neural Control
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
Queue Length Regret Bounds for Contextual Queueing Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems
von: Kang, Yue, et al.
Veröffentlicht: (2024) -
Online Continuous Hyperparameter Optimization for Generalized Linear Contextual Bandits
von: Kang, Yue, et al.
Veröffentlicht: (2023) -
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
von: Ye, Chenlu, et al.
Veröffentlicht: (2025) -
Generalized Low-Rank Matrix Contextual Bandits with Graph Information
von: Wang, Yao, et al.
Veröffentlicht: (2025) -
Lipschitz Bandits with Stochastic Delayed Feedback
von: Liu, Zhongxuan, et al.
Veröffentlicht: (2025)