RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
Fuente:
arXiv
Saved in:
| Main Authors: | Reber, David, Richardson, Sean, Nief, Todd, Garbacea, Cristina, Veitch, Victor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
Why is constrained neural language generation particularly challenging?
by: Garbacea, Cristina, et al.
Published: (2022)
by: Garbacea, Cristina, et al.
Published: (2022)
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
by: Gui, Lin, et al.
Published: (2024)
by: Gui, Lin, et al.
Published: (2024)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
by: Muchane, Mark, et al.
Published: (2025)
by: Muchane, Mark, et al.
Published: (2025)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026)
by: Garbacea, Cristina, et al.
Published: (2026)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
by: Vashishtha, Aniket, et al.
Published: (2023)
by: Vashishtha, Aniket, et al.
Published: (2023)
Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
by: Nief, Todd, et al.
Published: (2025)
by: Nief, Todd, et al.
Published: (2025)
Evaluating the Goal-Directedness of Large Language Models
by: Everitt, Tom, et al.
Published: (2025)
by: Everitt, Tom, et al.
Published: (2025)
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals
by: Zheng, Haoran, et al.
Published: (2024)
by: Zheng, Haoran, et al.
Published: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023)
by: Park, Kiho, et al.
Published: (2023)
Reward Models Identify Consistency, Not Causality
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
by: Shinoda, Kazutoshi, et al.
Published: (2026)
by: Shinoda, Kazutoshi, et al.
Published: (2026)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
by: Park, Kiho, et al.
Published: (2024)
by: Park, Kiho, et al.
Published: (2024)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
by: Zheng, Congmin, et al.
Published: (2025)
by: Zheng, Congmin, et al.
Published: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
by: Elle
Published: (2025)
by: Elle
Published: (2025)
Event Causality Identification with Synthetic Control
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Causal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies
by: Zeng, Donghuo, et al.
Published: (2025)
by: Zeng, Donghuo, et al.
Published: (2025)
Generative Framework for Personalized Persuasion: Inferring Causal, Counterfactual, and Latent Knowledge
by: Zeng, Donghuo, et al.
Published: (2025)
by: Zeng, Donghuo, et al.
Published: (2025)
Take its Essence, Discard its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect
by: Lu, Junyu, et al.
Published: (2024)
by: Lu, Junyu, et al.
Published: (2024)
Aligning Large Language Models with Counterfactual DPO
by: Butcher, Bradley
Published: (2024)
by: Butcher, Bradley
Published: (2024)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models
by: Apartsin, Alexander, et al.
Published: (2026)
by: Apartsin, Alexander, et al.
Published: (2026)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Tiny Reward Models
by: Pan, Sarah
Published: (2025)
by: Pan, Sarah
Published: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
CLOMO: Counterfactual Logical Modification with Large Language Models
by: Huang, Yinya, et al.
Published: (2023)
by: Huang, Yinya, et al.
Published: (2023)
HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
by: Garbacea, Cristina, et al.
Published: (2025)
by: Garbacea, Cristina, et al.
Published: (2025)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
by: Peng, Hao, et al.
Published: (2026)
by: Peng, Hao, et al.
Published: (2026)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
by: Da, Jeff, et al.
Published: (2025)
by: Da, Jeff, et al.
Published: (2025)
Self-Rewarding Language Models
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
by: Feng, Yijun
Published: (2025)
by: Feng, Yijun
Published: (2025)
Shadow-Loom: Causal Reasoning over Graphical World Models of Narratives
by: Wilmot, David
Published: (2026)
by: Wilmot, David
Published: (2026)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
by: Giorgi, Flavio, et al.
Published: (2024)
by: Giorgi, Flavio, et al.
Published: (2024)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
by: Kim, Jongho, et al.
Published: (2025)
by: Kim, Jongho, et al.
Published: (2025)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026)
by: Merrill, Scott, et al.
Published: (2026)
Similar Items
-
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026) -
Why is constrained neural language generation particularly challenging?
by: Garbacea, Cristina, et al.
Published: (2022) -
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
by: Gui, Lin, et al.
Published: (2024) -
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
by: Muchane, Mark, et al.
Published: (2025) -
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026)