When Can Proxies Improve the Sample Complexity of Preference Learning?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yuchen, de Souza, Daniel Augusto, Shi, Zhengyan, Yang, Mengyue, Minervini, Pasquale, D'Amour, Alexander, Kusner, Matt J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Choosing a Proxy Metric from Past Experiments
by: Tripuraneni, Nilesh, et al.
Published: (2023)
by: Tripuraneni, Nilesh, et al.
Published: (2023)
An Auditing Test To Detect Behavioral Shift in Language Models
by: Richter, Leo, et al.
Published: (2024)
by: Richter, Leo, et al.
Published: (2024)
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
by: Zheng, Jiajing, et al.
Published: (2021)
by: Zheng, Jiajing, et al.
Published: (2021)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
Deconfounding Scores and Representation Learning for Causal Effect Estimation with Weak Overlap
by: Clivio, Oscar, et al.
Published: (2026)
by: Clivio, Oscar, et al.
Published: (2026)
Mind the Graph When Balancing Data for Fairness or Robustness
by: Schrouff, Jessica, et al.
Published: (2024)
by: Schrouff, Jessica, et al.
Published: (2024)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
by: Lum, Kristian, et al.
Published: (2024)
by: Lum, Kristian, et al.
Published: (2024)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Making physical activity fun and accessible to adults with intellectual disabilities: A pilot study of a gamification intervention
by: Stéphanie Turgeon, et al.
Published: (2024)
by: Stéphanie Turgeon, et al.
Published: (2024)
Theoretical guarantees on the best-of-n alignment policy
by: Beirami, Ahmad, et al.
Published: (2024)
by: Beirami, Ahmad, et al.
Published: (2024)
Can LLMs Correct Physicians, Yet? Investigating Effective Interaction Methods in the Medical Domain
by: Sayin, Burcu, et al.
Published: (2024)
by: Sayin, Burcu, et al.
Published: (2024)
Predictive Churn with the Set of Good Models
by: Watson-Daniels, Jamelle, et al.
Published: (2024)
by: Watson-Daniels, Jamelle, et al.
Published: (2024)
Using Natural Language Explanations to Improve Robustness of In-context Learning
by: He, Xuanli, et al.
Published: (2023)
by: He, Xuanli, et al.
Published: (2023)
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Is Complex Query Answering Really Complex?
by: Gregucci, Cosimo, et al.
Published: (2024)
by: Gregucci, Cosimo, et al.
Published: (2024)
Expected Reward Prediction, with Applications to Model Routing
by: Hasanaliyev, Kenan, et al.
Published: (2026)
by: Hasanaliyev, Kenan, et al.
Published: (2026)
Causal Machine Learning: A Survey and Open Problems
by: Kaddour, Jean, et al.
Published: (2022)
by: Kaddour, Jean, et al.
Published: (2022)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
by: Saxena, Rohit, et al.
Published: (2026)
by: Saxena, Rohit, et al.
Published: (2026)
Low-Rank Compression of Language Models via Differentiable Rank Selection
by: Sundrani, Sidhant, et al.
Published: (2025)
by: Sundrani, Sidhant, et al.
Published: (2025)
Answerability in Retrieval-Augmented Open-Domain Question Answering
by: Abdumalikov, Rustam, et al.
Published: (2024)
by: Abdumalikov, Rustam, et al.
Published: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Process-Supervised Multi-Agent Reinforcement Learning for Reliable Clinical Reasoning
by: Lee, Chaeeun, et al.
Published: (2026)
by: Lee, Chaeeun, et al.
Published: (2026)
Setting the Record Straight on Transformer Oversmoothing
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
by: Yu, Simon, et al.
Published: (2024)
by: Yu, Simon, et al.
Published: (2024)
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
by: Shi, Zhengyan
Published: (2025)
by: Shi, Zhengyan
Published: (2025)
On the index of minimal 2-tori in the 4-sphere
by: Kusner, Rob, et al.
Published: (2018)
by: Kusner, Rob, et al.
Published: (2018)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Probing the Emergence of Cross-lingual Alignment during LLM Training
by: Wang, Hetong, et al.
Published: (2024)
by: Wang, Hetong, et al.
Published: (2024)
On the Independence Assumption in Neurosymbolic Learning
by: van Krieken, Emile, et al.
Published: (2024)
by: van Krieken, Emile, et al.
Published: (2024)
HPS: Hard Preference Sampling for Human Preference Alignment
by: Zou, Xiandong, et al.
Published: (2025)
by: Zou, Xiandong, et al.
Published: (2025)
Vibration Sensitivity of one-port and two-port MEMS microphones
by: Doyon-D'Amour, Francis, et al.
Published: (2024)
by: Doyon-D'Amour, Francis, et al.
Published: (2024)
The Sacketts [VHS - video] / / Escritor: Louis L´Amor; Director: Robert Totten
by: L´Amour, Louis
by: L´Amour, Louis
End of the drive / Louis L´Amour
by: L´Amour, Louis
by: L´Amour, Louis
Temporal Smoothness Regularisers for Neural Link Predictors
by: Dileo, Manuel, et al.
Published: (2023)
by: Dileo, Manuel, et al.
Published: (2023)
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
by: Devoto, Alessio, et al.
Published: (2024)
by: Devoto, Alessio, et al.
Published: (2024)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
by: Ghazaryan, Gayane, et al.
Published: (2024)
by: Ghazaryan, Gayane, et al.
Published: (2024)
Controlled expansion for transport in a class of non-Fermi liquids
by: Shi, Zhengyan Darius
Published: (2023)
by: Shi, Zhengyan Darius
Published: (2023)
Learning to Solve Complex Problems via Dataset Decomposition
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
Similar Items
-
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024) -
Choosing a Proxy Metric from Past Experiments
by: Tripuraneni, Nilesh, et al.
Published: (2023) -
An Auditing Test To Detect Behavioral Shift in Language Models
by: Richter, Leo, et al.
Published: (2024) -
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
by: Zheng, Jiajing, et al.
Published: (2021) -
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)