Soft Best-of-n Sampling for Model Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Verdun, Claudio Mayrink, Oesterling, Alex, Lakkaraju, Himabindu, Calmon, Flavio P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
Multi-Group Proportional Representation in Retrieval
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025)
by: Khalaf, Hadi, et al.
Published: (2025)
Optimized Couplings for Watermarking Large Language Models
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Multi-Group Proportional Representation for Text-to-Image Models
by: Jung, Sangwon, et al.
Published: (2025)
by: Jung, Sangwon, et al.
Published: (2025)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)
by: Oesterling, Alex, et al.
Published: (2023)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024)
by: Kumar, Aounon, et al.
Published: (2024)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
AI Alignment at Your Discretion
by: Buyl, Maarten, et al.
Published: (2025)
by: Buyl, Maarten, et al.
Published: (2025)
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
by: Gilani, Atefeh, et al.
Published: (2026)
by: Gilani, Atefeh, et al.
Published: (2026)
ProofCompass: Enhancing Specialized Provers with LLM Guidance
by: Wischermann, Nicolas, et al.
Published: (2025)
by: Wischermann, Nicolas, et al.
Published: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
High-Dimensional Confidence Regions in Sparse MRI
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Non-Asymptotic Uncertainty Quantification in High-Dimensional Learning
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Fixed-Budget Differentially Private Best Arm Identification
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
by: Chen, Zhirui, et al.
Published: (2025)
by: Chen, Zhirui, et al.
Published: (2025)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
Best Agent Identification for General Game Playing
by: Stephenson, Matthew, et al.
Published: (2025)
by: Stephenson, Matthew, et al.
Published: (2025)
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
by: Xiong, Zidi, et al.
Published: (2025)
by: Xiong, Zidi, et al.
Published: (2025)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
Understanding Learning through the Lens of Dynamical Invariants
by: Ushveridze, Alex
Published: (2024)
by: Ushveridze, Alex
Published: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
Federated Latent Space Alignment for Multi-user Semantic Communications
by: Di Poce, Giuseppe, et al.
Published: (2026)
by: Di Poce, Giuseppe, et al.
Published: (2026)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
With or Without Replacement? Improving Confidence in Fourier Imaging
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
by: Bian, Jie, et al.
Published: (2025)
by: Bian, Jie, et al.
Published: (2025)
The Alignment Bottleneck
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
by: Hou, Yunlong, et al.
Published: (2024)
by: Hou, Yunlong, et al.
Published: (2024)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
by: Badrinath, Charumathi, et al.
Published: (2024)
by: Badrinath, Charumathi, et al.
Published: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Towards Interpretable Soft Prompts
by: Patel, Oam, et al.
Published: (2025)
by: Patel, Oam, et al.
Published: (2025)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Deep Learning-based Compressive Beam Alignment in mmWave Vehicular Systems
by: Wang, Yuyang, et al.
Published: (2021)
by: Wang, Yuyang, et al.
Published: (2021)
Order-Optimal Sample Complexity of Rectified Flows
by: Sahoo, Hari Krishna, et al.
Published: (2026)
by: Sahoo, Hari Krishna, et al.
Published: (2026)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Approximate Implication for Probabilistic Graphical Models
by: Kenig, Batya
Published: (2023)
by: Kenig, Batya
Published: (2023)
Large AI Models for Wireless Physical Layer
by: Guo, Jiajia, et al.
Published: (2025)
by: Guo, Jiajia, et al.
Published: (2025)
Similar Items
-
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025) -
Multi-Group Proportional Representation in Retrieval
by: Oesterling, Alex, et al.
Published: (2024) -
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025) -
Optimized Couplings for Watermarking Large Language Models
by: Tsur, Dor, et al.
Published: (2025) -
Multi-Group Proportional Representation for Text-to-Image Models
by: Jung, Sangwon, et al.
Published: (2025)