Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhiwei, Liu, Hui, Li, Xiaomin, Dai, Zhenwei, Zeng, Jingying, Wang, Fali, Lin, Minhua, Chandradevan, Ramraj, Li, Zhen, Luo, Chen, Tang, Xianfeng, He, Qi, Wang, Suhang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
por: Zhang, Zhiwei, et al.
Publicado: (2025)
por: Zhang, Zhiwei, et al.
Publicado: (2025)
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
por: Wang, Fali, et al.
Publicado: (2025)
por: Wang, Fali, et al.
Publicado: (2025)
How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
por: Lin, Minhua, et al.
Publicado: (2026)
por: Lin, Minhua, et al.
Publicado: (2026)
How Far are LLMs from Real Search? A Comprehensive Study on Efficiency, Completeness, and Inherent Capabilities
por: Lin, Minhua, et al.
Publicado: (2025)
por: Lin, Minhua, et al.
Publicado: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
por: Zhang, Zhiwei, et al.
Publicado: (2024)
por: Zhang, Zhiwei, et al.
Publicado: (2024)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
por: Cui, Yingqian, et al.
Publicado: (2025)
por: Cui, Yingqian, et al.
Publicado: (2025)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
por: Zeng, Jingying, et al.
Publicado: (2025)
por: Zeng, Jingying, et al.
Publicado: (2025)
A General Framework to Enhance Fine-tuning-based LLM Unlearning
por: Ren, Jie, et al.
Publicado: (2025)
por: Ren, Jie, et al.
Publicado: (2025)
Generative Query Reformulation Using Ensemble Prompting, Document Fusion, and Relevance Feedback
por: Dhole, Kaustubh D., et al.
Publicado: (2024)
por: Dhole, Kaustubh D., et al.
Publicado: (2024)
DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation
por: Chandradevan, Ramraj, et al.
Publicado: (2024)
por: Chandradevan, Ramraj, et al.
Publicado: (2024)
Examples as the Prompt: A Scalable Approach for Efficient LLM Adaptation in E-Commerce
por: Zeng, Jingying, et al.
Publicado: (2025)
por: Zeng, Jingying, et al.
Publicado: (2025)
MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution
por: Lin, Minhua, et al.
Publicado: (2026)
por: Lin, Minhua, et al.
Publicado: (2026)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
por: Cui, Yingqian, et al.
Publicado: (2025)
por: Cui, Yingqian, et al.
Publicado: (2025)
A Survey of Calibration Process for Black-Box LLMs
por: Xie, Liangru, et al.
Publicado: (2024)
por: Xie, Liangru, et al.
Publicado: (2024)
To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems
por: He, Pengfei, et al.
Publicado: (2025)
por: He, Pengfei, et al.
Publicado: (2025)
Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
por: Wu, Zongyu, et al.
Publicado: (2025)
por: Wu, Zongyu, et al.
Publicado: (2025)
Rethinking Graph Backdoor Attacks: A Distribution-Preserving Perspective
por: Zhang, Zhiwei, et al.
Publicado: (2024)
por: Zhang, Zhiwei, et al.
Publicado: (2024)
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
por: Wang, Fali, et al.
Publicado: (2025)
por: Wang, Fali, et al.
Publicado: (2025)
QueryExplorer: An Interactive Query Generation Assistant for Search and Exploration
por: Dhole, Kaustubh D., et al.
Publicado: (2024)
por: Dhole, Kaustubh D., et al.
Publicado: (2024)
Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
por: Ren, Jie, et al.
Publicado: (2025)
por: Ren, Jie, et al.
Publicado: (2025)
PreGIP: Watermarking the Pretraining of Graph Neural Networks for Deep Intellectual Property Protection
por: Dai, Enyan, et al.
Publicado: (2024)
por: Dai, Enyan, et al.
Publicado: (2024)
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
por: Wang, Fali, et al.
Publicado: (2024)
por: Wang, Fali, et al.
Publicado: (2024)
LLM and GNN are Complementary: Distilling LLM for Multimodal Graph Learning
por: Xu, Junjie, et al.
Publicado: (2024)
por: Xu, Junjie, et al.
Publicado: (2024)
Bradley-Terry Policy Optimization for Generative Preference Modeling
por: Feng, Shengyu, et al.
Publicado: (2025)
por: Feng, Shengyu, et al.
Publicado: (2025)
PageRank and the Bradley-Terry model
por: Selby, David Antony
Publicado: (2024)
por: Selby, David Antony
Publicado: (2024)
The Bradley-Terry Stochastic Block Model
por: Santi, Lapo, et al.
Publicado: (2025)
por: Santi, Lapo, et al.
Publicado: (2025)
Efficient Inference for Covariate-adjusted Bradley-Terry Model with Covariate Shift
por: Li, Xiudi, et al.
Publicado: (2025)
por: Li, Xiudi, et al.
Publicado: (2025)
Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
Recent advances in the Bradley--Terry Model: theory, algorithms, and applications
por: Fang, Shuxing, et al.
Publicado: (2026)
por: Fang, Shuxing, et al.
Publicado: (2026)
The many routes to the ubiquitous Bradley-Terry model
por: Hamilton, Ian, et al.
Publicado: (2023)
por: Hamilton, Ian, et al.
Publicado: (2023)
Position: Agentic Evolution is the Path to Evolving LLMs
por: Lin, Minhua, et al.
Publicado: (2026)
por: Lin, Minhua, et al.
Publicado: (2026)
A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
por: Lin, Minhua, et al.
Publicado: (2025)
por: Lin, Minhua, et al.
Publicado: (2025)
A spectral approach for the dynamic Bradley–Terry model
por: Xinyu Tian, et al.
Publicado: (2024)
por: Xinyu Tian, et al.
Publicado: (2024)
Minimax Hypothesis Testing for the Bradley-Terry-Luce Model
por: Makur, Anuran, et al.
Publicado: (2024)
por: Makur, Anuran, et al.
Publicado: (2024)
Robustness Inspired Graph Backdoor Defense
por: Zhang, Zhiwei, et al.
Publicado: (2024)
por: Zhang, Zhiwei, et al.
Publicado: (2024)
HC-GST: Heterophily-aware Distribution Consistency based Graph Self-training
por: Wang, Fali, et al.
Publicado: (2024)
por: Wang, Fali, et al.
Publicado: (2024)
A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
por: Wang, Fali, et al.
Publicado: (2024)
por: Wang, Fali, et al.
Publicado: (2024)
Neural Bradley-Terry Rating: Quantifying Properties from Comparisons
por: Fujii, Satoru
Publicado: (2023)
por: Fujii, Satoru
Publicado: (2023)
Are You Using Reliable Graph Prompts? Trojan Prompt Attacks on Graph Neural Networks
por: Lin, Minhua, et al.
Publicado: (2024)
por: Lin, Minhua, et al.
Publicado: (2024)
GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning
por: Wang, Fali, et al.
Publicado: (2026)
por: Wang, Fali, et al.
Publicado: (2026)
Ejemplares similares
-
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
por: Zhang, Zhiwei, et al.
Publicado: (2025) -
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
por: Wang, Fali, et al.
Publicado: (2025) -
How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
por: Lin, Minhua, et al.
Publicado: (2026) -
How Far are LLMs from Real Search? A Comprehensive Study on Efficiency, Completeness, and Inherent Capabilities
por: Lin, Minhua, et al.
Publicado: (2025) -
Catastrophic Failure of LLM Unlearning via Quantization
por: Zhang, Zhiwei, et al.
Publicado: (2024)