Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
Fuente:
arXiv
Saved in:
| Main Authors: | Rashidinejad, Paria, Tian, Yuandong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026)
by: Mustafi, Aratrika, et al.
Published: (2026)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026)
by: Nguyen, Minh
Published: (2026)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
by: Bareilles, Gilles, et al.
Published: (2023)
by: Bareilles, Gilles, et al.
Published: (2023)
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
by: Mohamed, Mimoun, et al.
Published: (2023)
by: Mohamed, Mimoun, et al.
Published: (2023)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
by: Kitaoka, Akira
Published: (2025)
by: Kitaoka, Akira
Published: (2025)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
by: He, Yixuan, et al.
Published: (2025)
by: He, Yixuan, et al.
Published: (2025)
FraPPE: Fast and Efficient Preference-based Pure Exploration
by: Das, Udvas, et al.
Published: (2025)
by: Das, Udvas, et al.
Published: (2025)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
by: Yan, Shunxing, et al.
Published: (2026)
by: Yan, Shunxing, et al.
Published: (2026)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Reward-Directed Score-Based Diffusion Models via q-Learning
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Robustly Learning Monotone Generalized Linear Models via Data Augmentation
by: Zarifis, Nikos, et al.
Published: (2025)
by: Zarifis, Nikos, et al.
Published: (2025)
Robustly Learning Single-Index Models via Alignment Sharpness
by: Zarifis, Nikos, et al.
Published: (2024)
by: Zarifis, Nikos, et al.
Published: (2024)
Robust stochastic first order methods in heavy-tailed noise via medoid mini-batch gradient sampling
by: Vukovic, Manojlo, et al.
Published: (2026)
by: Vukovic, Manojlo, et al.
Published: (2026)
Learning the Uncertainty Sets for Control Dynamics via Set Membership: A Non-Asymptotic Analysis
by: Li, Yingying, et al.
Published: (2023)
by: Li, Yingying, et al.
Published: (2023)
Robust Assortment Optimization from Observational Data
by: Lu, Miao, et al.
Published: (2026)
by: Lu, Miao, et al.
Published: (2026)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Failure of uniform laws of large numbers for subdifferentials and beyond
by: Tian, Lai, et al.
Published: (2025)
by: Tian, Lai, et al.
Published: (2025)
Geometry-induced Regularization in Deep ReLU Neural Networks
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Inference-Time Alignment for Diffusion Models via Variationally Stable Doob's Matching
by: Chang, Jinyuan, et al.
Published: (2026)
by: Chang, Jinyuan, et al.
Published: (2026)
Blessings and Curses of Covariate Shifts: Adversarial Learning Dynamics, Directional Convergence, and Equilibria
by: Liang, Tengyuan
Published: (2022)
by: Liang, Tengyuan
Published: (2022)
Data-Efficient Non-Gaussian Semi-Nonparametric Density Estimation for Nonlinear Dynamical Systems
by: Liao, Aaron R., et al.
Published: (2026)
by: Liao, Aaron R., et al.
Published: (2026)
On Regularization via Early Stopping for Least Squares Regression
by: Sonthalia, Rishi, et al.
Published: (2024)
by: Sonthalia, Rishi, et al.
Published: (2024)
Distributionally Robust Instrumental Variables Estimation
by: Qu, Zhaonan, et al.
Published: (2024)
by: Qu, Zhaonan, et al.
Published: (2024)
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
by: Karnik, Santhosh, et al.
Published: (2024)
by: Karnik, Santhosh, et al.
Published: (2024)
Algorithms for mean-field variational inference via polyhedral optimization in the Wasserstein space
by: Jiang, Yiheng, et al.
Published: (2023)
by: Jiang, Yiheng, et al.
Published: (2023)
Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences
by: Aolaritei, Liviu, et al.
Published: (2025)
by: Aolaritei, Liviu, et al.
Published: (2025)
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
by: Roy, Somjit, et al.
Published: (2026)
by: Roy, Somjit, et al.
Published: (2026)
Hyperparameter tuning via trajectory predictions: Stochastic prox-linear methods in matrix sensing
by: Lou, Mengqi, et al.
Published: (2024)
by: Lou, Mengqi, et al.
Published: (2024)
Joint learning of a network of linear dynamical systems via total variation penalization
by: Donnat, Claire, et al.
Published: (2025)
by: Donnat, Claire, et al.
Published: (2025)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
by: Cheng, Xiuyuan, et al.
Published: (2023)
by: Cheng, Xiuyuan, et al.
Published: (2023)
Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees
by: Maros, Marie, et al.
Published: (2022)
by: Maros, Marie, et al.
Published: (2022)
Joint Learning of Linear Dynamical Systems under Smoothness Constraints
by: Tyagi, Hemant
Published: (2024)
by: Tyagi, Hemant
Published: (2024)
Similar Items
-
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026) -
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024) -
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026) -
Piecewise Polynomial Regression of Tame Functions via Integer Programming
by: Bareilles, Gilles, et al.
Published: (2023) -
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)