Debiasing Reward Models by Representation Learning with Guarantees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ng, Ignavier, Blöbaum, Patrick, Bhandari, Siddharth, Zhang, Kun, Kasiviswanathan, Shiva |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Anytime-Valid Inference for Double/Debiased Machine Learning of Causal Parameters
von: Dalal, Abhinandan, et al.
Veröffentlicht: (2024)
von: Dalal, Abhinandan, et al.
Veröffentlicht: (2024)
Learning to Answer from Correct Demonstrations
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
von: Hiremath, Sujai, et al.
Veröffentlicht: (2025)
von: Hiremath, Sujai, et al.
Veröffentlicht: (2025)
Continual Learning of Nonlinear Independent Representations
von: Sun, Boyang, et al.
Veröffentlicht: (2024)
von: Sun, Boyang, et al.
Veröffentlicht: (2024)
On the Identifiability of Nonlinear ICA: Sparsity and Beyond
von: Zheng, Yujia, et al.
Veröffentlicht: (2022)
von: Zheng, Yujia, et al.
Veröffentlicht: (2022)
A Classical View on Benign Overfitting: The Role of Sample Size
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
von: Park, Junhyung, et al.
Veröffentlicht: (2024)
von: Park, Junhyung, et al.
Veröffentlicht: (2024)
Modeling Causal Mechanisms with Diffusion Models for Interventional and Counterfactual Queries
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
Local Causal Discovery with Linear non-Gaussian Cyclic Models
von: Dai, Haoyue, et al.
Veröffentlicht: (2024)
von: Dai, Haoyue, et al.
Veröffentlicht: (2024)
A Quantitative Characterization of Forgetting in Post-Training
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
Sequential Kernelized Independence Testing
von: Podkopaev, Aleksandr, et al.
Veröffentlicht: (2022)
von: Podkopaev, Aleksandr, et al.
Veröffentlicht: (2022)
What Causes Postoperative Aspiration?
von: Nagesh, Supriya, et al.
Veröffentlicht: (2025)
von: Nagesh, Supriya, et al.
Veröffentlicht: (2025)
Federated Causal Discovery from Heterogeneous Data
von: Li, Loka, et al.
Veröffentlicht: (2024)
von: Li, Loka, et al.
Veröffentlicht: (2024)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
Debiased Model-based Representations for Sample-efficient Continuous Control
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
Causal Representation Learning from Multiple Distributions: A General Setting
von: Zhang, Kun, et al.
Veröffentlicht: (2024)
von: Zhang, Kun, et al.
Veröffentlicht: (2024)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
Autoencoding Conditional Neural Processes for Representation Learning
von: Prokhorov, Victor, et al.
Veröffentlicht: (2023)
von: Prokhorov, Victor, et al.
Veröffentlicht: (2023)
The PetShop Dataset -- Finding Causes of Performance Issues across Microservices
von: Hardt, Michaela, et al.
Veröffentlicht: (2023)
von: Hardt, Michaela, et al.
Veröffentlicht: (2023)
Structure Learning with Continuous Optimization: A Sober Look and Beyond
von: Ng, Ignavier, et al.
Veröffentlicht: (2023)
von: Ng, Ignavier, et al.
Veröffentlicht: (2023)
Modeling Feature Maps for Quantum Machine Learning
von: Singh, Navneet, et al.
Veröffentlicht: (2025)
von: Singh, Navneet, et al.
Veröffentlicht: (2025)
Limits of Approximating the Median Treatment Effect
von: Addanki, Raghavendra, et al.
Veröffentlicht: (2024)
von: Addanki, Raghavendra, et al.
Veröffentlicht: (2024)
Causal vs. Anticausal merging of predictors
von: Mejia, Sergio Hernan Garrido, et al.
Veröffentlicht: (2025)
von: Mejia, Sergio Hernan Garrido, et al.
Veröffentlicht: (2025)
Auxiliary Reward Generation with Transition Distance Representation Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
Debiasing Machine Unlearning with Counterfactual Examples
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal models
von: Blöbaum, Patrick, et al.
Veröffentlicht: (2022)
von: Blöbaum, Patrick, et al.
Veröffentlicht: (2022)
Debiasing Kernel-Based Generative Models
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
Causal Representation Learning from General Environments under Nonparametric Mixing
von: Ng, Ignavier, et al.
Veröffentlicht: (2026)
von: Ng, Ignavier, et al.
Veröffentlicht: (2026)
ReLA: Representation Learning and Aggregation for Job Scheduling with Reinforcement Learning
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2026)
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2026)
Notes on the Reward Representation of Posterior Updates
von: Ortega, Pedro A.
Veröffentlicht: (2026)
von: Ortega, Pedro A.
Veröffentlicht: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
von: Mohamed, Faisal, et al.
Veröffentlicht: (2026)
von: Mohamed, Faisal, et al.
Veröffentlicht: (2026)
Debiasing Online Preference Learning via Preference Feature Preservation
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
An Independent Implementation of Quantum Machine Learning Algorithms in Qiskit for Genomic Data
von: Singh, Navneet, et al.
Veröffentlicht: (2024)
von: Singh, Navneet, et al.
Veröffentlicht: (2024)
Reward Models in Deep Reinforcement Learning: A Survey
von: Yu, Rui, et al.
Veröffentlicht: (2025)
von: Yu, Rui, et al.
Veröffentlicht: (2025)
Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling
von: He, Shenghong
Veröffentlicht: (2025)
von: He, Shenghong
Veröffentlicht: (2025)
Inference Time Debiasing Concepts in Diffusion Models
von: Kupssinskü, Lucas S., et al.
Veröffentlicht: (2025)
von: Kupssinskü, Lucas S., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Anytime-Valid Inference for Double/Debiased Machine Learning of Causal Parameters
von: Dalal, Abhinandan, et al.
Veröffentlicht: (2024) -
Learning to Answer from Correct Demonstrations
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025) -
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
von: Hiremath, Sujai, et al.
Veröffentlicht: (2025) -
Continual Learning of Nonlinear Independent Representations
von: Sun, Boyang, et al.
Veröffentlicht: (2024) -
On the Identifiability of Nonlinear ICA: Sparsity and Beyond
von: Zheng, Yujia, et al.
Veröffentlicht: (2022)