Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yupei, Yang, Lin, Deng, Wanxi, Qu, Lin, Feng, Fan, Huang, Biwei, Tu, Shikui, Xu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations
by: Yang, Yupei, et al.
Published: (2024)
by: Yang, Yupei, et al.
Published: (2024)
Boosting Efficiency in Task-Agnostic Exploration through Causal Knowledge
by: Yang, Yupei, et al.
Published: (2024)
by: Yang, Yupei, et al.
Published: (2024)
Hallucination-Resistant Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
by: Yang, Yupei, et al.
Published: (2025)
by: Yang, Yupei, et al.
Published: (2025)
UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates
by: Yang, Yupei, et al.
Published: (2026)
by: Yang, Yupei, et al.
Published: (2026)
Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
by: Lin, Guang, et al.
Published: (2026)
by: Lin, Guang, et al.
Published: (2026)
Learning a Pessimistic Reward Model in RLHF
by: Xu, Yinglun, et al.
Published: (2025)
by: Xu, Yinglun, et al.
Published: (2025)
Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow Networks
by: Qian, Hao, et al.
Published: (2025)
by: Qian, Hao, et al.
Published: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Causal Representation Meets Stochastic Modeling under Generic Geometry
by: Ren, Jiaxu, et al.
Published: (2026)
by: Ren, Jiaxu, et al.
Published: (2026)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
by: Xu, Wenyuan, et al.
Published: (2025)
by: Xu, Wenyuan, et al.
Published: (2025)
How to Evaluate Reward Models for RLHF
by: Frick, Evan, et al.
Published: (2024)
by: Frick, Evan, et al.
Published: (2024)
Causal Structure Learning in Hawkes Processes with Complex Latent Confounder Networks
by: Jin, Songyao, et al.
Published: (2025)
by: Jin, Songyao, et al.
Published: (2025)
THFlow: A Temporally Hierarchical Flow Matching Framework for 3D Peptide Design
by: Huang, Dengdeng, et al.
Published: (2025)
by: Huang, Dengdeng, et al.
Published: (2025)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom Number
by: Zhou, Jingyuan, et al.
Published: (2025)
by: Zhou, Jingyuan, et al.
Published: (2025)
Optimal Design for Reward Modeling in RLHF
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Differentiable Causal Discovery For Latent Hierarchical Causal Models
by: Prashant, Parjanya, et al.
Published: (2024)
by: Prashant, Parjanya, et al.
Published: (2024)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Robust Post-Training for Generative Recommenders: Why Exponential Reward-Weighted SFT Outperforms RLHF
by: Chidambaram, Keertana, et al.
Published: (2026)
by: Chidambaram, Keertana, et al.
Published: (2026)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Generative design and validation of therapeutic peptides for glioblastoma based on a potential target ATP5A
by: Qian, Hao, et al.
Published: (2025)
by: Qian, Hao, et al.
Published: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
Reinforcement Learning for Causal Discovery without Acyclicity Constraints
by: Duong, Bao, et al.
Published: (2024)
by: Duong, Bao, et al.
Published: (2024)
Reward Generalization in RLHF: A Topological Perspective
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Robust Reward Modeling via Causal Rubrics
by: Srivastava, Pragya, et al.
Published: (2025)
by: Srivastava, Pragya, et al.
Published: (2025)
A Fast Kernel-based Conditional Independence test with Application to Causal Discovery
by: Schacht, Oliver, et al.
Published: (2025)
by: Schacht, Oliver, et al.
Published: (2025)
Quantile Regression for Distributional Reward Models in RLHF
by: Dorka, Nicolai
Published: (2024)
by: Dorka, Nicolai
Published: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
by: Tang, Jinzhou, et al.
Published: (2026)
by: Tang, Jinzhou, et al.
Published: (2026)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
by: Liu, Jing
Published: (2025)
by: Liu, Jing
Published: (2025)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
Information-Theoretic Reward Decomposition for Generalizable RLHF
by: Mao, Liyuan, et al.
Published: (2025)
by: Mao, Liyuan, et al.
Published: (2025)
Transformer Is Inherently a Causal Learner
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Learning Identifiable Factorized Causal Representations of Cellular Responses
by: Mao, Haiyi, et al.
Published: (2024)
by: Mao, Haiyi, et al.
Published: (2024)
Similar Items
-
Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations
by: Yang, Yupei, et al.
Published: (2024) -
Boosting Efficiency in Task-Agnostic Exploration through Causal Knowledge
by: Yang, Yupei, et al.
Published: (2024) -
Hallucination-Resistant Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
by: Yang, Yupei, et al.
Published: (2025) -
UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates
by: Yang, Yupei, et al.
Published: (2026) -
Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
by: Lin, Guang, et al.
Published: (2026)