Bayesian Reward Models for LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Adam X., Robeyns, Maxime, Coste, Thomas, Shi, Zhengyan, Wang, Jun, Bou-Ammar, Haitham, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Improving LLM-Generated Code Quality with GRPO
by: Robeyns, Maxime, et al.
Published: (2025)
by: Robeyns, Maxime, et al.
Published: (2025)
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
by: Mower, Christopher E., et al.
Published: (2025)
by: Mower, Christopher E., et al.
Published: (2025)
Contextual Causal Bayesian Optimisation
by: Arsenyan, Vahan, et al.
Published: (2023)
by: Arsenyan, Vahan, et al.
Published: (2023)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
A Self-Improving Coding Agent
by: Robeyns, Maxime, et al.
Published: (2025)
by: Robeyns, Maxime, et al.
Published: (2025)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
by: Ji, Xiaotong, et al.
Published: (2025)
by: Ji, Xiaotong, et al.
Published: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
by: Tutnov, Rasul, et al.
Published: (2025)
by: Tutnov, Rasul, et al.
Published: (2025)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
by: Oomerjee, Adnan, et al.
Published: (2025)
by: Oomerjee, Adnan, et al.
Published: (2025)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Efficient Reinforcement Learning with Large Language Model Priors
by: Yan, Xue, et al.
Published: (2024)
by: Yan, Xue, et al.
Published: (2024)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
by: Hazard, Hugo, et al.
Published: (2025)
by: Hazard, Hugo, et al.
Published: (2025)
Safe Reinforcement Learning on the Constraint Manifold: Theory and Applications
by: Liu, Puze, et al.
Published: (2024)
by: Liu, Puze, et al.
Published: (2024)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
by: Fountas, Zafeirios, et al.
Published: (2026)
by: Fountas, Zafeirios, et al.
Published: (2026)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
by: Nguyen, Tu, et al.
Published: (2026)
by: Nguyen, Tu, et al.
Published: (2026)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025)
by: Wieser, Frederico, et al.
Published: (2025)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Reward Model Ensembles Help Mitigate Overoptimization
by: Coste, Thomas, et al.
Published: (2023)
by: Coste, Thomas, et al.
Published: (2023)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
by: Benfeghoul, Martin, et al.
Published: (2025)
by: Benfeghoul, Martin, et al.
Published: (2025)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
by: Bou, Matthieu, et al.
Published: (2025)
by: Bou, Matthieu, et al.
Published: (2025)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Using Neural Networks for Data Cleaning in Weather Datasets
by: Hanslope, Jack R. P., et al.
Published: (2024)
by: Hanslope, Jack R. P., et al.
Published: (2024)
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024)
by: Milsom, Edward, et al.
Published: (2024)
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023)
by: Milsom, Edward, et al.
Published: (2023)
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025)
by: Milsom, Edward, et al.
Published: (2025)
Similar Items
-
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023) -
Improving LLM-Generated Code Quality with GRPO
by: Robeyns, Maxime, et al.
Published: (2025) -
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023) -
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
by: Mower, Christopher E., et al.
Published: (2025) -
Contextual Causal Bayesian Optimisation
by: Arsenyan, Vahan, et al.
Published: (2023)