Scalable Ensembling For Mitigating Reward Overoptimisation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ahmed, Ahmed M., Rafailov, Rafael, Sharkov, Stepan, Li, Xuechen, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
Extracting books from production language models
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026)
von: Chen, Edward, et al.
Veröffentlicht: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
General Preference Reinforcement Learning
von: Umer, Muhammad, et al.
Veröffentlicht: (2026)
von: Umer, Muhammad, et al.
Veröffentlicht: (2026)
Why Do Safety Guardrails Degrade Across Languages?
von: Zhang, Max, et al.
Veröffentlicht: (2026)
von: Zhang, Max, et al.
Veröffentlicht: (2026)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
Scaling Laws for Downstream Task Performance of Large Language Models
von: Isik, Berivan, et al.
Veröffentlicht: (2024)
von: Isik, Berivan, et al.
Veröffentlicht: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
Disentangling Length from Quality in Direct Preference Optimization
von: Park, Ryan, et al.
Veröffentlicht: (2024)
von: Park, Ryan, et al.
Veröffentlicht: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
von: Chan, Willy, et al.
Veröffentlicht: (2025)
von: Chan, Willy, et al.
Veröffentlicht: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2025)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2025)
Is Pre-training Truly Better Than Meta-Learning?
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
von: Chawla, Krrish, et al.
Veröffentlicht: (2025)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
von: Lu, Keming, et al.
Veröffentlicht: (2024)
von: Lu, Keming, et al.
Veröffentlicht: (2024)
BanglaASTE: A Novel Framework for Aspect-Sentiment-Opinion Extraction in Bangla E-commerce Reviews Using Ensemble Deep Learning
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
Causally Inspired Regularization Enables Domain General Representations
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Scalable Multi-phase Word Embedding Using Conjunctive Propositional Clauses
von: Kadhim, Ahmed K., et al.
Veröffentlicht: (2025)
von: Kadhim, Ahmed K., et al.
Veröffentlicht: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
von: Rho, Donghwan
Veröffentlicht: (2025)
von: Rho, Donghwan
Veröffentlicht: (2025)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025) -
Extracting books from production language models
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026) -
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026) -
Reasoning Models Don't Just Think Longer, They Move Differently
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026) -
General Preference Reinforcement Learning
von: Umer, Muhammad, et al.
Veröffentlicht: (2026)