A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
Fuente:
arXiv
Saved in:
| Main Authors: | LeVine, Will, Pikus, Benjamin, Chen, Anthony, Hendryx, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Out-of-Distribution Detection & Applications With Ablated Learned Temperature Energy
by: LeVine, Will, et al.
Published: (2024)
by: LeVine, Will, et al.
Published: (2024)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025)
by: Gunjal, Anisha, et al.
Published: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
Foundation Model's Embedded Representations May Detect Distribution Shift
by: Vargas, Max, et al.
Published: (2023)
by: Vargas, Max, et al.
Published: (2023)
Omitted Variable Bias in Language Models Under Distribution Shift
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
by: Nath, Vaskar, et al.
Published: (2025)
by: Nath, Vaskar, et al.
Published: (2025)
Generalizing Reward Modeling for Out-of-Distribution Preference Learning
by: Jia, Chen
Published: (2024)
by: Jia, Chen
Published: (2024)
On-Policy RL with Optimal Reward Baseline
by: Hao, Yaru, et al.
Published: (2025)
by: Hao, Yaru, et al.
Published: (2025)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
by: Li, Yufei, et al.
Published: (2024)
by: Li, Yufei, et al.
Published: (2024)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
by: Young, Sean I.
Published: (2024)
by: Young, Sean I.
Published: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
by: Das, Debeshee, et al.
Published: (2024)
by: Das, Debeshee, et al.
Published: (2024)
Evaluating the Process Modeling Abilities of Large Language Models -- Preliminary Foundations and Results
by: Fettke, Peter, et al.
Published: (2025)
by: Fettke, Peter, et al.
Published: (2025)
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
by: Yu, Le, et al.
Published: (2023)
by: Yu, Le, et al.
Published: (2023)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
Quantile Regression for Distributional Reward Models in RLHF
by: Dorka, Nicolai
Published: (2024)
by: Dorka, Nicolai
Published: (2024)
Text Classification Under Class Distribution Shift: A Survey
by: Costache, Adriana Valentina, et al.
Published: (2025)
by: Costache, Adriana Valentina, et al.
Published: (2025)
Continual Learning Under Language Shift
by: Gogoulou, Evangelia, et al.
Published: (2023)
by: Gogoulou, Evangelia, et al.
Published: (2023)
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
by: Sedova, Anastasiia, et al.
Published: (2024)
by: Sedova, Anastasiia, et al.
Published: (2024)
Battery powered. The promise of energy storage / Steve LeVine
by: LeVine, Steve
Published: (2014)
by: LeVine, Steve
Published: (2014)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024)
by: Muckatira, Sherin, et al.
Published: (2024)
Leveraging AI Graders for Missing Score Imputation to Achieve Accurate Ability Estimation in Constructed-Response Tests
by: Uto, Masaki, et al.
Published: (2025)
by: Uto, Masaki, et al.
Published: (2025)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
by: Ma, Qiyao, et al.
Published: (2026)
by: Ma, Qiyao, et al.
Published: (2026)
Noise Contrastive Alignment of Language Models with Explicit Rewards
by: Chen, Huayu, et al.
Published: (2024)
by: Chen, Huayu, et al.
Published: (2024)
Entropy-Regularized Process Reward Model
by: Zhang, Hanning, et al.
Published: (2024)
by: Zhang, Hanning, et al.
Published: (2024)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)
by: Mathur, Leena, et al.
Published: (2025)
A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
by: Tuan, Yi-Lin, et al.
Published: (2024)
by: Tuan, Yi-Lin, et al.
Published: (2024)
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Expected Reward Prediction, with Applications to Model Routing
by: Hasanaliyev, Kenan, et al.
Published: (2026)
by: Hasanaliyev, Kenan, et al.
Published: (2026)
Learning Goal-Conditioned Representations for Language Reward Models
by: Nath, Vaskar, et al.
Published: (2024)
by: Nath, Vaskar, et al.
Published: (2024)
Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts
by: Choi, Jihye, et al.
Published: (2024)
by: Choi, Jihye, et al.
Published: (2024)
Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
by: Zhao, Zekai, et al.
Published: (2025)
by: Zhao, Zekai, et al.
Published: (2025)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Similar Items
-
Out-of-Distribution Detection & Applications With Ablated Learned Temperature Energy
by: LeVine, Will, et al.
Published: (2024) -
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025) -
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026) -
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)