Theoretical guarantees on the best-of-n alignment policy
Fuente:
arXiv
Saved in:
| Main Authors: | Beirami, Ahmad, Agarwal, Alekh, Berant, Jonathan, D'Amour, Alexander, Eisenstein, Jacob, Nagpal, Chirag, Suresh, Ananda Theertha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024)
by: Balashankar, Ananth, et al.
Published: (2024)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
by: Setlur, Amrith, et al.
Published: (2024)
by: Setlur, Amrith, et al.
Published: (2024)
SpecTr: Fast Speculative Decoding via Optimal Transport
by: Sun, Ziteng, et al.
Published: (2023)
by: Sun, Ziteng, et al.
Published: (2023)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
by: Lum, Kristian, et al.
Published: (2024)
by: Lum, Kristian, et al.
Published: (2024)
Block Verification Accelerates Speculative Decoding
by: Sun, Ziteng, et al.
Published: (2024)
by: Sun, Ziteng, et al.
Published: (2024)
Asymptotics of Language Model Alignment
by: Yang, Joy Qiping, et al.
Published: (2024)
by: Yang, Joy Qiping, et al.
Published: (2024)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
by: Paes, Lucas Monteiro, et al.
Published: (2023)
by: Paes, Lucas Monteiro, et al.
Published: (2023)
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
Cost-Optimal Active AI Model Evaluation
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
On Robust Hypothesis Testing with respect to the Hellinger Distance
by: Modak, Eeshan, et al.
Published: (2025)
by: Modak, Eeshan, et al.
Published: (2025)
Mean estimation in the add-remove model of differential privacy
by: Kulesza, Alex, et al.
Published: (2023)
by: Kulesza, Alex, et al.
Published: (2023)
Coupling without Communication and Drafter-Invariant Speculative Decoding
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
The importance of feature preprocessing for differentially private linear optimization
by: Sun, Ziteng, et al.
Published: (2023)
by: Sun, Ziteng, et al.
Published: (2023)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
by: Wu, Zhaofeng, et al.
Published: (2024)
by: Wu, Zhaofeng, et al.
Published: (2024)
Rate of Model Collapse in Recursive Training
by: Suresh, Ananda Theertha, et al.
Published: (2024)
by: Suresh, Ananda Theertha, et al.
Published: (2024)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
by: Eisenstein, Jacob, et al.
Published: (2026)
by: Eisenstein, Jacob, et al.
Published: (2026)
Subset-Based Instance Optimality in Private Estimation
by: Dick, Travis, et al.
Published: (2023)
by: Dick, Travis, et al.
Published: (2023)
ALTA: Compiler-Based Analysis of Transformers
by: Shaw, Peter, et al.
Published: (2024)
by: Shaw, Peter, et al.
Published: (2024)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
by: Zheng, Jiajing, et al.
Published: (2021)
by: Zheng, Jiajing, et al.
Published: (2021)
ARAGOG: Advanced RAG Output Grading
by: Eibich, Matouš, et al.
Published: (2024)
by: Eibich, Matouš, et al.
Published: (2024)
Improving Robustness via Tilted Exponential Layer: A Communication-Theoretic Perspective
by: Puranik, Bhagyashree, et al.
Published: (2023)
by: Puranik, Bhagyashree, et al.
Published: (2023)
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval
by: Rubin, Ohad, et al.
Published: (2023)
by: Rubin, Ohad, et al.
Published: (2023)
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
by: Ye, Haotian, et al.
Published: (2025)
by: Ye, Haotian, et al.
Published: (2025)
Don't lie to your friends: Learning what you know from collaborative self-play
by: Eisenstein, Jacob, et al.
Published: (2025)
by: Eisenstein, Jacob, et al.
Published: (2025)
Graph Fusion Across Languages using Large Language Models
by: Kyaw, Kaung Myat, et al.
Published: (2026)
by: Kyaw, Kaung Myat, et al.
Published: (2026)
Expected Reward Prediction, with Applications to Model Routing
by: Hasanaliyev, Kenan, et al.
Published: (2026)
by: Hasanaliyev, Kenan, et al.
Published: (2026)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
by: Marinov, Teodor V., et al.
Published: (2024)
by: Marinov, Teodor V., et al.
Published: (2024)
Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
by: Amos, Ido, et al.
Published: (2023)
by: Amos, Ido, et al.
Published: (2023)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Weakly Supervised Text-to-SQL Parsing through Question Decomposition
by: Wolfson, Tomer, et al.
Published: (2021)
by: Wolfson, Tomer, et al.
Published: (2021)
Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker
by: Saxena, Rachna, et al.
Published: (2025)
by: Saxena, Rachna, et al.
Published: (2025)
Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering Datasets
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2025)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2025)
A Representation Sharpening Framework for Zero Shot Dense Retrieval
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Preference Models assume Proportional Hazards of Utilities
by: Nagpal, Chirag
Published: (2025)
by: Nagpal, Chirag
Published: (2025)
Similar Items
-
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024) -
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024) -
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024) -
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023) -
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
by: Setlur, Amrith, et al.
Published: (2024)