Robust Reward Alignment via Hypothesis Space Batch Cutting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Zhixian, Zhang, Haode, Feng, Yizhe, Jin, Wanxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback
von: Vatsa, Amitesh, et al.
Veröffentlicht: (2026)
von: Vatsa, Amitesh, et al.
Veröffentlicht: (2026)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
Test-Time Alignment via Hypothesis Reweighting
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
von: Mahesheka, Harsh, et al.
Veröffentlicht: (2024)
von: Mahesheka, Harsh, et al.
Veröffentlicht: (2024)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
von: Vergara-Browne, Tomás, et al.
Veröffentlicht: (2026)
von: Vergara-Browne, Tomás, et al.
Veröffentlicht: (2026)
Energy-Based Reward Models for Robust Language Model Alignment
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
Robust Batched Bandits
von: Guo, Yunwen, et al.
Veröffentlicht: (2025)
von: Guo, Yunwen, et al.
Veröffentlicht: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling
von: Lou, Xingzhou, et al.
Veröffentlicht: (2024)
von: Lou, Xingzhou, et al.
Veröffentlicht: (2024)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
von: Li, Yanshi, et al.
Veröffentlicht: (2024)
von: Li, Yanshi, et al.
Veröffentlicht: (2024)
Robust Reward Modeling via Causal Rubrics
von: Srivastava, Pragya, et al.
Veröffentlicht: (2025)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2025)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
COMBA: Cross Batch Aggregation for Learning Large Graphs with Context Gating State Space Models
von: Shen, Jiajun, et al.
Veröffentlicht: (2026)
von: Shen, Jiajun, et al.
Veröffentlicht: (2026)
Non-Convex Robust Hypothesis Testing using Sinkhorn Uncertainty Sets
von: Wang, Jie, et al.
Veröffentlicht: (2024)
von: Wang, Jie, et al.
Veröffentlicht: (2024)
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation
von: Liu, Runze, et al.
Veröffentlicht: (2023)
von: Liu, Runze, et al.
Veröffentlicht: (2023)
Bayesian Reward Models for LLM Alignment
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
Hypothesis Spaces for Deep Learning
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
Batch Active Learning of Reward Functions from Human Preferences
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
High-Dimensional Robust Mean Estimation with Untrusted Batches
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2026)
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2026)
Improving Group Robustness on Spurious Correlation via Evidential Alignment
von: Ye, Wenqian, et al.
Veröffentlicht: (2025)
von: Ye, Wenqian, et al.
Veröffentlicht: (2025)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Latent Space Translation via Semantic Alignment
von: Maiorca, Valentino, et al.
Veröffentlicht: (2023)
von: Maiorca, Valentino, et al.
Veröffentlicht: (2023)
Closed-Form Concept Erasure via Double Projections
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
Diversified Batch Selection for Training Acceleration
von: Hong, Feng, et al.
Veröffentlicht: (2024)
von: Hong, Feng, et al.
Veröffentlicht: (2024)
Distributed Differentiable Dynamic Game for Multi-robot Coordination
von: Zhou, Yizhi, et al.
Veröffentlicht: (2022)
von: Zhou, Yizhi, et al.
Veröffentlicht: (2022)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
Federated-inspired Single-cell Batch Integration in Latent Space
von: Nguyen, Quang-Huy, et al.
Veröffentlicht: (2026)
von: Nguyen, Quang-Huy, et al.
Veröffentlicht: (2026)
Beyond Manual Annotation: A Human-AI Collaborative Framework for Medical Image Segmentation Using Only "Better or Worse" Expert Feedback
von: Zhang, Yizhe
Veröffentlicht: (2025)
von: Zhang, Yizhe
Veröffentlicht: (2025)
High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch
von: Liautaud, Paul, et al.
Veröffentlicht: (2026)
von: Liautaud, Paul, et al.
Veröffentlicht: (2026)
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2024)
Latent Space Communication via K-V Cache Alignment
von: Dery, Lucio M., et al.
Veröffentlicht: (2026)
von: Dery, Lucio M., et al.
Veröffentlicht: (2026)
GETA-3DGS: Automatic Joint Structured Pruning and Quantization for 3D Gaussian Splatting
von: Zhang, Baobing, et al.
Veröffentlicht: (2026)
von: Zhang, Baobing, et al.
Veröffentlicht: (2026)
Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward
von: Jia, Zhiwei, et al.
Veröffentlicht: (2024)
von: Jia, Zhiwei, et al.
Veröffentlicht: (2024)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback
von: Vatsa, Amitesh, et al.
Veröffentlicht: (2026) -
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
von: Sahu, Sharan, et al.
Veröffentlicht: (2025) -
Test-Time Alignment via Hypothesis Reweighting
von: Lee, Yoonho, et al.
Veröffentlicht: (2024) -
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
von: Mahesheka, Harsh, et al.
Veröffentlicht: (2024) -
Operationalising the Superficial Alignment Hypothesis via Task Complexity
von: Vergara-Browne, Tomás, et al.
Veröffentlicht: (2026)