Interpreting Language Reward Models via Contrastive Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Junqi, Bewley, Tom, Mishra, Saumitra, Lecue, Freddy, Veloso, Manuela |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
by: Hedström, Anna, et al.
Published: (2025)
by: Hedström, Anna, et al.
Published: (2025)
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
by: Kariyappa, Sanjay, et al.
Published: (2024)
by: Kariyappa, Sanjay, et al.
Published: (2024)
ShapShift: Explaining Model Prediction Shifts with Subgroup Conditional Shapley Values
by: Bewley, Tom, et al.
Published: (2026)
by: Bewley, Tom, et al.
Published: (2026)
Entropic Projection Alignment: Estimating, Explaining, and Improving Model Performance Under Distribution Shift
by: Amoukou, Salim I., et al.
Published: (2026)
by: Amoukou, Salim I., et al.
Published: (2026)
Sequential Harmful Shift Detection Without Labels
by: Amoukou, Salim I., et al.
Published: (2024)
by: Amoukou, Salim I., et al.
Published: (2024)
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
by: Hamman, Faisal, et al.
Published: (2024)
by: Hamman, Faisal, et al.
Published: (2024)
The Effect of Data Poisoning on Counterfactual Explanations
by: Artelt, André, et al.
Published: (2024)
by: Artelt, André, et al.
Published: (2024)
Counterfactual Metarules for Local and Global Recourse
by: Bewley, Tom, et al.
Published: (2024)
by: Bewley, Tom, et al.
Published: (2024)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
by: Hamman, Faisal, et al.
Published: (2023)
by: Hamman, Faisal, et al.
Published: (2023)
Interval Abstractions for Robust Counterfactual Explanations
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
Provably Robust and Plausible Counterfactual Explanations for Neural Networks via Robust Optimisation
by: Jiang, Junqi, et al.
Published: (2023)
by: Jiang, Junqi, et al.
Published: (2023)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
by: Liu, Dengcan, et al.
Published: (2026)
by: Liu, Dengcan, et al.
Published: (2026)
Robust Counterfactual Explanations in Machine Learning: A Survey
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
RobustX: Robust Counterfactual Explanations Made Easy
by: Jiang, Junqi, et al.
Published: (2025)
by: Jiang, Junqi, et al.
Published: (2025)
Zero-Shot Reinforcement Learning from Low Quality Data
by: Jeen, Scott, et al.
Published: (2023)
by: Jeen, Scott, et al.
Published: (2023)
Zero-Shot Reinforcement Learning Under Partial Observability
by: Jeen, Scott, et al.
Published: (2025)
by: Jeen, Scott, et al.
Published: (2025)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
by: Jiang, Junqi, et al.
Published: (2025)
by: Jiang, Junqi, et al.
Published: (2025)
A Visual Tool for Interactive Model Explanation using Sensitivity Analysis
by: Schuler, Manuela
Published: (2025)
by: Schuler, Manuela
Published: (2025)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
by: Bui, Ngoc, et al.
Published: (2025)
by: Bui, Ngoc, et al.
Published: (2025)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed Samples
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
Interpreting Inflammation Prediction Model via Tag-based Cohort Explanation
by: Meng, Fanyu, et al.
Published: (2024)
by: Meng, Fanyu, et al.
Published: (2024)
CELL your Model: Contrastive Explanations for Large Language Models
by: Luss, Ronny, et al.
Published: (2024)
by: Luss, Ronny, et al.
Published: (2024)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Recourse under Model Multiplicity via Argumentative Ensembling (Technical Report)
by: Jiang, Junqi, et al.
Published: (2023)
by: Jiang, Junqi, et al.
Published: (2023)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Causality-Aware Local Interpretable Model-Agnostic Explanations
by: Cinquini, Martina, et al.
Published: (2022)
by: Cinquini, Martina, et al.
Published: (2022)
Explanation-Guided Adversarial Training for Robust and Interpretable Models
by: Chen, Chao, et al.
Published: (2026)
by: Chen, Chao, et al.
Published: (2026)
Concept-Based Abductive and Contrastive Explanations for Behaviors of Vision Models
by: Canizales, Ronaldo, et al.
Published: (2026)
by: Canizales, Ronaldo, et al.
Published: (2026)
Learning Mathematical Rules with Large Language Models
by: Gorceix, Antoine, et al.
Published: (2024)
by: Gorceix, Antoine, et al.
Published: (2024)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Efficient Process Reward Modeling via Contrastive Mutual Information
by: Lee, Nakyung, et al.
Published: (2026)
by: Lee, Nakyung, et al.
Published: (2026)
Global Concept Explanations for Graphs by Contrastive Learning
by: Teufel, Jonas, et al.
Published: (2024)
by: Teufel, Jonas, et al.
Published: (2024)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
by: Maghsoudi, Maryam, et al.
Published: (2026)
by: Maghsoudi, Maryam, et al.
Published: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
Are Logistic Models Really Interpretable?
by: Dervovic, Danial, et al.
Published: (2024)
by: Dervovic, Danial, et al.
Published: (2024)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
by: Zhang, Feng, et al.
Published: (2026)
by: Zhang, Feng, et al.
Published: (2026)
TACENR: Task-Agnostic Contrastive Explanations for Node Representations
by: Papanikou, Vasiliki, et al.
Published: (2026)
by: Papanikou, Vasiliki, et al.
Published: (2026)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Similar Items
-
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
by: Hedström, Anna, et al.
Published: (2025) -
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
by: Kariyappa, Sanjay, et al.
Published: (2024) -
ShapShift: Explaining Model Prediction Shifts with Subgroup Conditional Shapley Values
by: Bewley, Tom, et al.
Published: (2026) -
Entropic Projection Alignment: Estimating, Explaining, and Improving Model Performance Under Distribution Shift
by: Amoukou, Salim I., et al.
Published: (2026) -
Sequential Harmful Shift Detection Without Labels
by: Amoukou, Salim I., et al.
Published: (2024)