Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
Fuente:
arXiv
Saved in:
| Main Authors: | Lamparth, Max, Fein, Daniel, Haupt, Andreas, Hussing, Marcel, Kochenderfer, Mykel J. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
by: Fein, Daniel, et al.
Published: (2026)
by: Fein, Daniel, et al.
Published: (2026)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025)
by: Chaubard, Francois, et al.
Published: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)
by: Zhao, Kangwen, et al.
Published: (2025)
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)
by: Chaubard, Francois, et al.
Published: (2024)
Conditional Deep Generative Models for Belief State Planning
by: Bigeard, Antoine, et al.
Published: (2025)
by: Bigeard, Antoine, et al.
Published: (2025)
Graph Q-Learning for Combinatorial Optimization
by: Dax, Victoria M., et al.
Published: (2024)
by: Dax, Victoria M., et al.
Published: (2024)
Robust Planning for Autonomous Vehicles with Diffusion-Based Failure Samplers
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
A Semi-Decentralized Approach to Multiagent Control
by: Al-Husseini, Mahdi, et al.
Published: (2026)
by: Al-Husseini, Mahdi, et al.
Published: (2026)
Addressing Myopic Constrained POMDP Planning with Recursive Dual Ascent
by: Stocco, Paula, et al.
Published: (2024)
by: Stocco, Paula, et al.
Published: (2024)
Semi-Markovian Planning to Coordinate Aerial and Maritime Medical Evacuation Platforms
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
by: Pan, Jinhao, et al.
Published: (2026)
by: Pan, Jinhao, et al.
Published: (2026)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
by: Jafari, Kiana, et al.
Published: (2026)
by: Jafari, Kiana, et al.
Published: (2026)
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
by: Eddy, Duncan, et al.
Published: (2024)
by: Eddy, Duncan, et al.
Published: (2024)
Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments
by: Lange, Bernard, et al.
Published: (2023)
by: Lange, Bernard, et al.
Published: (2023)
BetaZero: Belief-State Planning for Long-Horizon POMDPs using Learned Approximations
by: Moss, Robert J., et al.
Published: (2023)
by: Moss, Robert J., et al.
Published: (2023)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
On Technique Identification and Threat-Actor Attribution using LLMs and Embedding Models
by: Guru, Kyla, et al.
Published: (2025)
by: Guru, Kyla, et al.
Published: (2025)
Fair Learning for Bias Mitigation and Quality Optimization in Paper Recommendation
by: Oyshi, Uttamasha Anjally, et al.
Published: (2026)
by: Oyshi, Uttamasha Anjally, et al.
Published: (2026)
CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
by: Blankenstein, Thierry, et al.
Published: (2025)
by: Blankenstein, Thierry, et al.
Published: (2025)
Diffusion Models for Safety Validation of Autonomous Driving Systems
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
Large-Scale Multi-Robot Assembly Planning for Autonomous Manufacturing
by: Brown, Kyle, et al.
Published: (2023)
by: Brown, Kyle, et al.
Published: (2023)
A New Strategy for Verifying Reach-Avoid Specifications in Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
ConstrainedZero: Chance-Constrained POMDP Planning using Learned Probabilistic Failure Surrogates and Adaptive Safety Constraints
by: Moss, Robert J., et al.
Published: (2024)
by: Moss, Robert J., et al.
Published: (2024)
Zono-Conformal Prediction: Zonotope-Based Uncertainty Quantification for Regression and Classification Tasks
by: Lützow, Laura, et al.
Published: (2025)
by: Lützow, Laura, et al.
Published: (2025)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
by: Shrivastava, Aryan, et al.
Published: (2024)
by: Shrivastava, Aryan, et al.
Published: (2024)
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
by: Grabb, Declan, et al.
Published: (2024)
by: Grabb, Declan, et al.
Published: (2024)
Mitigating Cognitive Bias in RLHF by Altering Rationality
by: Horter, Tiffany, et al.
Published: (2026)
by: Horter, Tiffany, et al.
Published: (2026)
Importance Sampling-Guided Meta-Training for Intelligent Agents in Highly Interactive Environments
by: Arief, Mansur, et al.
Published: (2024)
by: Arief, Mansur, et al.
Published: (2024)
Improving the Resilience of Quadrotors in Underground Environments by Combining Learning-based and Safety Controllers
by: Ward, Isaac Ronald, et al.
Published: (2025)
by: Ward, Isaac Ronald, et al.
Published: (2025)
Subgroups Matter for Robust Bias Mitigation
by: Alloula, Anissa, et al.
Published: (2025)
by: Alloula, Anissa, et al.
Published: (2025)
The FABRIC Strategy for Verifying Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
Mitigating Selection Bias with Node Pruning and Auxiliary Options
by: Choi, Hyeong Kyu, et al.
Published: (2024)
by: Choi, Hyeong Kyu, et al.
Published: (2024)
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
by: Dunivin, Zackary Okun, et al.
Published: (2026)
by: Dunivin, Zackary Okun, et al.
Published: (2026)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
by: Zhou, Hanzhang, et al.
Published: (2024)
by: Zhou, Hanzhang, et al.
Published: (2024)
Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
by: Hickman, Louis, et al.
Published: (2024)
by: Hickman, Louis, et al.
Published: (2024)
Similar Items
-
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
by: Fein, Daniel, et al.
Published: (2026) -
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024) -
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025) -
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025) -
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)