Bradley-Terry Policy Optimization for Generative Preference Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Feng, Shengyu, He, Yun, Ma, Shuang, Li, Beibin, Xiong, Yuanhao, Li, Songlin, Mandyam, Karishma, Katz-Samuels, Julian, Bi, Shengjie, Yu, Licheng, Zhang, Hejia, Sankararaman, Karthik Abinav, Fang, Han, Yang, Yiming, Faruqui, Manaal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
por: He, Yun, et al.
Publicado: (2025)
por: He, Yun, et al.
Publicado: (2025)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
Recent advances in the Bradley--Terry Model: theory, algorithms, and applications
por: Fang, Shuxing, et al.
Publicado: (2026)
por: Fang, Shuxing, et al.
Publicado: (2026)
Contextual Bandits with Packing and Covering Constraints: A Modular Lagrangian Approach via Regression
por: Slivkins, Aleksandrs, et al.
Publicado: (2022)
por: Slivkins, Aleksandrs, et al.
Publicado: (2022)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
por: He, Yun, et al.
Publicado: (2024)
por: He, Yun, et al.
Publicado: (2024)
PageRank and the Bradley-Terry model
por: Selby, David Antony
Publicado: (2024)
por: Selby, David Antony
Publicado: (2024)
The Bradley-Terry Stochastic Block Model
por: Santi, Lapo, et al.
Publicado: (2025)
por: Santi, Lapo, et al.
Publicado: (2025)
Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model
por: Ge, Yurun, et al.
Publicado: (2025)
por: Ge, Yurun, et al.
Publicado: (2025)
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
por: Zhang, Zhiwei, et al.
Publicado: (2025)
por: Zhang, Zhiwei, et al.
Publicado: (2025)
Efficient Inference for Covariate-adjusted Bradley-Terry Model with Covariate Shift
por: Li, Xiudi, et al.
Publicado: (2025)
por: Li, Xiudi, et al.
Publicado: (2025)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
por: Hong, Yuzhong, et al.
Publicado: (2024)
por: Hong, Yuzhong, et al.
Publicado: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
The many routes to the ubiquitous Bradley-Terry model
por: Hamilton, Ian, et al.
Publicado: (2023)
por: Hamilton, Ian, et al.
Publicado: (2023)
A spectral approach for the dynamic Bradley–Terry model
por: Xinyu Tian, et al.
Publicado: (2024)
por: Xinyu Tian, et al.
Publicado: (2024)
Minimax Hypothesis Testing for the Bradley-Terry-Luce Model
por: Makur, Anuran, et al.
Publicado: (2024)
por: Makur, Anuran, et al.
Publicado: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
por: Xu, Tengyu, et al.
Publicado: (2024)
por: Xu, Tengyu, et al.
Publicado: (2024)
Neural Bradley-Terry Rating: Quantifying Properties from Comparisons
por: Fujii, Satoru
Publicado: (2023)
por: Fujii, Satoru
Publicado: (2023)
Preference Optimization with Multi-Sample Comparisons
por: Wang, Chaoqi, et al.
Publicado: (2024)
por: Wang, Chaoqi, et al.
Publicado: (2024)
To stay discovered: On tournament mean score sequences and the Bradley--Terry model
por: Aldous, David, et al.
Publicado: (2018)
por: Aldous, David, et al.
Publicado: (2018)
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation
por: Zhou, Shang, et al.
Publicado: (2026)
por: Zhou, Shang, et al.
Publicado: (2026)
The Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory
por: Okahara, Hisaya, et al.
Publicado: (2026)
por: Okahara, Hisaya, et al.
Publicado: (2026)
Generalized Bradley-Terry Models for Score Estimation from Paired Comparisons
por: Fageot, Julien, et al.
Publicado: (2023)
por: Fageot, Julien, et al.
Publicado: (2023)
The incomplete Analytic Hierarchy Process and Bradley-Terry model: (in)consistency and information retrieval
por: Gyarmati, László, et al.
Publicado: (2022)
por: Gyarmati, László, et al.
Publicado: (2022)
Inclusive Ranking of Indian States and Union Territories via Bayesian Bradley-Terry Model
por: Rizvi, Arshi, et al.
Publicado: (2026)
por: Rizvi, Arshi, et al.
Publicado: (2026)
Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach
por: Ocampo, Nicolás Benjamín, et al.
Publicado: (2026)
por: Ocampo, Nicolás Benjamín, et al.
Publicado: (2026)
Scalable Bayesian Inference for Bradley--Terry Models with Ties: An Application to Honour Based Abuse
por: Seymour, Rowland G, et al.
Publicado: (2024)
por: Seymour, Rowland G, et al.
Publicado: (2024)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
por: Vu, Tu, et al.
Publicado: (2024)
por: Vu, Tu, et al.
Publicado: (2024)
Generalized Parallel Scaling with Interdependent Generations
por: Dong, Harry, et al.
Publicado: (2025)
por: Dong, Harry, et al.
Publicado: (2025)
HYPO: Hyperspherical Out-of-Distribution Generalization
por: Bai, Haoyue, et al.
Publicado: (2024)
por: Bai, Haoyue, et al.
Publicado: (2024)
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
por: Krishna, Satyapriya, et al.
Publicado: (2024)
por: Krishna, Satyapriya, et al.
Publicado: (2024)
What Matters for Model Merging at Scale?
por: Yadav, Prateek, et al.
Publicado: (2024)
por: Yadav, Prateek, et al.
Publicado: (2024)
Inference in a generalized Bradley-Terry model for paired comparisons with covariates and a growing number of subjects
por: Yan, Ting
Publicado: (2025)
por: Yan, Ting
Publicado: (2025)
An analysis of factors impacting team strengths in the Australian Football League using time-variant Bradley-Terry models
por: Soffner, Carlos Rafael Gonzalez, et al.
Publicado: (2024)
por: Soffner, Carlos Rafael Gonzalez, et al.
Publicado: (2024)
AutoMix: Automatically Mixing Language Models
por: Aggarwal, Pranjal, et al.
Publicado: (2023)
por: Aggarwal, Pranjal, et al.
Publicado: (2023)
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
por: Zhou, Runlong, et al.
Publicado: (2024)
por: Zhou, Runlong, et al.
Publicado: (2024)
HoSNN: Adversarially-Robust Homeostatic Spiking Neural Networks with Adaptive Firing Thresholds
por: Geng, Hejia, et al.
Publicado: (2023)
por: Geng, Hejia, et al.
Publicado: (2023)
Reinforcement Learning from User Feedback
por: Han, Eric, et al.
Publicado: (2025)
por: Han, Eric, et al.
Publicado: (2025)
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
por: Rao, Abinav, et al.
Publicado: (2026)
por: Rao, Abinav, et al.
Publicado: (2026)
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
por: Corrado, Nicholas E., et al.
Publicado: (2025)
por: Corrado, Nicholas E., et al.
Publicado: (2025)
Ejemplares similares
-
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
por: He, Yun, et al.
Publicado: (2025) -
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024) -
Recent advances in the Bradley--Terry Model: theory, algorithms, and applications
por: Fang, Shuxing, et al.
Publicado: (2026) -
Contextual Bandits with Packing and Covering Constraints: A Modular Lagrangian Approach via Regression
por: Slivkins, Aleksandrs, et al.
Publicado: (2022) -
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
por: He, Yun, et al.
Publicado: (2024)