Rectifying Shortcut Behaviors in Preference-based Reward Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Wenqian, Zheng, Guangtao, Zhang, Aidong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
Improving Group Robustness on Spurious Correlation via Evidential Alignment
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
Learning Robust Classifiers with Self-Guided Spurious Correlation Mitigation
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
Benchmarking Spurious Bias in Few-Shot Image Classifiers
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
A Comprehensive Survey on the Risks and Limitations of Concept-based Models
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
by: Sun, Shengjie, et al.
Published: (2025)
by: Sun, Shengjie, et al.
Published: (2025)
MiMu: Mitigating Multiple Shortcut Learning Behavior of Transformers
by: Zhao, Lili, et al.
Published: (2025)
by: Zhao, Lili, et al.
Published: (2025)
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
by: Choi, Heewoong, et al.
Published: (2024)
by: Choi, Heewoong, et al.
Published: (2024)
Residual Reward Models for Preference-based Reinforcement Learning
by: Cao, Chenyang, et al.
Published: (2025)
by: Cao, Chenyang, et al.
Published: (2025)
Reward Learning From Preference With Ties
by: Liu, Jinsong, et al.
Published: (2024)
by: Liu, Jinsong, et al.
Published: (2024)
In-Context Reward Adaptation for Robust Preference Modeling
by: Sun, Zhenyu, et al.
Published: (2026)
by: Sun, Zhenyu, et al.
Published: (2026)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
by: Metcalf, Katherine, et al.
Published: (2024)
by: Metcalf, Katherine, et al.
Published: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
CoLiDR: Concept Learning using Aggregated Disentangled Representations
by: Sinha, Sanchit, et al.
Published: (2024)
by: Sinha, Sanchit, et al.
Published: (2024)
ProtoNAM: Prototypical Neural Additive Models for Interpretable Deep Tabular Learning
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
by: Zhou, Zhiyuan, et al.
Published: (2022)
by: Zhou, Zhiyuan, et al.
Published: (2022)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Mitigating Shortcut Learning with InterpoLated Learning
by: Korakakis, Michalis, et al.
Published: (2025)
by: Korakakis, Michalis, et al.
Published: (2025)
Gradient-based Model Shortcut Detection for Time Series Classification
by: Ibarra, Salomon, et al.
Published: (2025)
by: Ibarra, Salomon, et al.
Published: (2025)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
Learning with Logical Constraints but without Shortcut Satisfaction
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Behavior Preference Regression for Offline Reinforcement Learning
by: Srinivasan, Padmanaba, et al.
Published: (2025)
by: Srinivasan, Padmanaba, et al.
Published: (2025)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Capturing Individual Human Preferences with Reward Features
by: Barreto, André, et al.
Published: (2025)
by: Barreto, André, et al.
Published: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
by: Ye, Ziyi, et al.
Published: (2024)
by: Ye, Ziyi, et al.
Published: (2024)
Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
A Typology for Exploring the Mitigation of Shortcut Behavior
by: Friedrich, Felix, et al.
Published: (2022)
by: Friedrich, Felix, et al.
Published: (2022)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
$i$REPO: $i$mplicit Reward Pairwise Difference based Empirical Preference Optimization
by: Le, Long Tan, et al.
Published: (2024)
by: Le, Long Tan, et al.
Published: (2024)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
by: Song, Haonan, et al.
Published: (2026)
by: Song, Haonan, et al.
Published: (2026)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
by: Pukdee, Rattana, et al.
Published: (2026)
by: Pukdee, Rattana, et al.
Published: (2026)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
Similar Items
-
ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
by: Zheng, Guangtao, et al.
Published: (2025) -
Improving Group Robustness on Spurious Correlation via Evidential Alignment
by: Ye, Wenqian, et al.
Published: (2025) -
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025) -
Learning Robust Classifiers with Self-Guided Spurious Correlation Mitigation
by: Zheng, Guangtao, et al.
Published: (2024) -
Benchmarking Spurious Bias in Few-Shot Image Classifiers
by: Zheng, Guangtao, et al.
Published: (2024)