Saved in:
| Main Authors: | Xie, Zhixian, Zhang, Haode, Feng, Yizhe, Jin, Wanxin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.02921 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback
by: Vatsa, Amitesh, et al.
Published: (2026)
by: Vatsa, Amitesh, et al.
Published: (2026)
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
by: Mahesheka, Harsh, et al.
Published: (2024)
by: Mahesheka, Harsh, et al.
Published: (2024)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
by: Sahu, Sharan, et al.
Published: (2025)
by: Sahu, Sharan, et al.
Published: (2025)
Test-Time Alignment via Hypothesis Reweighting
by: Lee, Yoonho, et al.
Published: (2024)
by: Lee, Yoonho, et al.
Published: (2024)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
by: Vergara-Browne, Tomás, et al.
Published: (2026)
by: Vergara-Browne, Tomás, et al.
Published: (2026)
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Robust Batched Bandits
by: Guo, Yunwen, et al.
Published: (2025)
by: Guo, Yunwen, et al.
Published: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
On the Robustness of Reward Models for Language Model Alignment
by: Hong, Jiwoo, et al.
Published: (2025)
by: Hong, Jiwoo, et al.
Published: (2025)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
by: Ye, Chenlu, et al.
Published: (2025)
by: Ye, Chenlu, et al.
Published: (2025)
SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
Distributed Differentiable Dynamic Game for Multi-robot Coordination
by: Zhou, Yizhi, et al.
Published: (2022)
by: Zhou, Yizhi, et al.
Published: (2022)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
by: Li, Yanshi, et al.
Published: (2024)
by: Li, Yanshi, et al.
Published: (2024)
Safe MPC Alignment with Human Directional Feedback
by: Xie, Zhixian, et al.
Published: (2024)
by: Xie, Zhixian, et al.
Published: (2024)
GETA-3DGS: Automatic Joint Structured Pruning and Quantization for 3D Gaussian Splatting
by: Zhang, Baobing, et al.
Published: (2026)
by: Zhang, Baobing, et al.
Published: (2026)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Closed-Form Concept Erasure via Double Projections
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Non-Convex Robust Hypothesis Testing using Sinkhorn Uncertainty Sets
by: Wang, Jie, et al.
Published: (2024)
by: Wang, Jie, et al.
Published: (2024)
COMBA: Cross Batch Aggregation for Learning Large Graphs with Context Gating State Space Models
by: Shen, Jiajun, et al.
Published: (2026)
by: Shen, Jiajun, et al.
Published: (2026)
Robust Reward Modeling via Causal Rubrics
by: Srivastava, Pragya, et al.
Published: (2025)
by: Srivastava, Pragya, et al.
Published: (2025)
Hypothesis Spaces for Deep Learning
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Beyond Manual Annotation: A Human-AI Collaborative Framework for Medical Image Segmentation Using Only "Better or Worse" Expert Feedback
by: Zhang, Yizhe
Published: (2025)
by: Zhang, Yizhe
Published: (2025)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Bayesian Reward Models for LLM Alignment
by: Yang, Adam X., et al.
Published: (2024)
by: Yang, Adam X., et al.
Published: (2024)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
by: Cheng, Tianhao, et al.
Published: (2026)
by: Cheng, Tianhao, et al.
Published: (2026)
PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation
by: Liu, Runze, et al.
Published: (2023)
by: Liu, Runze, et al.
Published: (2023)
Where to Touch, How to Contact: Hierarchical RL-MPC Framework for Geometry-Aware Long-Horizon Dexterous Manipulation
by: Xie, Zhixian, et al.
Published: (2026)
by: Xie, Zhixian, et al.
Published: (2026)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
High-Dimensional Robust Mean Estimation with Untrusted Batches
by: Aliakbarpour, Maryam, et al.
Published: (2026)
by: Aliakbarpour, Maryam, et al.
Published: (2026)
High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch
by: Liautaud, Paul, et al.
Published: (2026)
by: Liautaud, Paul, et al.
Published: (2026)
Unified Inference Framework for Single and Multi-Player Performative Prediction: Method and Asymptotic Optimality
by: Zhang, Zhixian, et al.
Published: (2026)
by: Zhang, Zhixian, et al.
Published: (2026)
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
Latent Space Communication via K-V Cache Alignment
by: Dery, Lucio M., et al.
Published: (2026)
by: Dery, Lucio M., et al.
Published: (2026)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
by: Fang, Linghan, et al.
Published: (2026)
by: Fang, Linghan, et al.
Published: (2026)
Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces
by: Zhu, Linglingzhi, et al.
Published: (2024)
by: Zhu, Linglingzhi, et al.
Published: (2024)
Similar Items
-
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback
by: Vatsa, Amitesh, et al.
Published: (2026) -
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
by: Mahesheka, Harsh, et al.
Published: (2024) -
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
by: Sahu, Sharan, et al.
Published: (2025) -
Test-Time Alignment via Hypothesis Reweighting
by: Lee, Yoonho, et al.
Published: (2024) -
Operationalising the Superficial Alignment Hypothesis via Task Complexity
by: Vergara-Browne, Tomás, et al.
Published: (2026)