The Perfect Blend: Redefining RLHF with Mixture of Judges
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Tengyu, Helenowski, Eryk, Sankararaman, Karthik Abinav, Jin, Di, Peng, Kaiyan, Han, Eric, Nie, Shaoliang, Zhu, Chen, Zhang, Hejia, Zhou, Wenxuan, Zeng, Zhouhao, He, Yun, Mandyam, Karishma, Talabzadeh, Arya, Khabsa, Madian, Cohen, Gabriel, Tian, Yuandong, Ma, Hao, Wang, Sinong, Fang, Han |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reinforcement Learning from User Feedback
por: Han, Eric, et al.
Publicado: (2025)
por: Han, Eric, et al.
Publicado: (2025)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
por: Yu, Zishun, et al.
Publicado: (2025)
por: Yu, Zishun, et al.
Publicado: (2025)
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
por: Qin, Chengwei, et al.
Publicado: (2025)
por: Qin, Chengwei, et al.
Publicado: (2025)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
por: He, Yun, et al.
Publicado: (2024)
por: He, Yun, et al.
Publicado: (2024)
Generalized Parallel Scaling with Interdependent Generations
por: Dong, Harry, et al.
Publicado: (2025)
por: Dong, Harry, et al.
Publicado: (2025)
Preference Optimization with Multi-Sample Comparisons
por: Wang, Chaoqi, et al.
Publicado: (2024)
por: Wang, Chaoqi, et al.
Publicado: (2024)
On the Equivalence of Graph Convolution and Mixup
por: Han, Xiaotian, et al.
Publicado: (2023)
por: Han, Xiaotian, et al.
Publicado: (2023)
Bradley-Terry Policy Optimization for Generative Preference Modeling
por: Feng, Shengyu, et al.
Publicado: (2025)
por: Feng, Shengyu, et al.
Publicado: (2025)
Contextual Bandits with Packing and Covering Constraints: A Modular Lagrangian Approach via Regression
por: Slivkins, Aleksandrs, et al.
Publicado: (2022)
por: Slivkins, Aleksandrs, et al.
Publicado: (2022)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
por: Yang, Xianjun, et al.
Publicado: (2025)
por: Yang, Xianjun, et al.
Publicado: (2025)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
por: Franzmeyer, Tim, et al.
Publicado: (2025)
por: Franzmeyer, Tim, et al.
Publicado: (2025)
Improving Model Factuality with Fine-grained Critique-based Evaluator
por: Xie, Yiqing, et al.
Publicado: (2024)
por: Xie, Yiqing, et al.
Publicado: (2024)
Improving Offline RL by Blending Heuristics
por: Geng, Sinong, et al.
Publicado: (2023)
por: Geng, Sinong, et al.
Publicado: (2023)
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
por: Xu, Zhiyang, et al.
Publicado: (2025)
por: Xu, Zhiyang, et al.
Publicado: (2025)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
por: Phan, Hoang, et al.
Publicado: (2025)
por: Phan, Hoang, et al.
Publicado: (2025)
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
por: Wu, Bo, et al.
Publicado: (2025)
por: Wu, Bo, et al.
Publicado: (2025)
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
por: Rao, Abinav, et al.
Publicado: (2026)
por: Rao, Abinav, et al.
Publicado: (2026)
Reuse and Blend: Energy-Efficient Optical Neural Network Enabled by Weight Sharing
por: Xu, Bo, et al.
Publicado: (2024)
por: Xu, Bo, et al.
Publicado: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Simulating, Visualizing and Playing with de Sitter and anti de Sitter spacetime
por: Kopczynski, Eryk
Publicado: (2023)
por: Kopczynski, Eryk
Publicado: (2023)
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
por: Bandarkar, Lucas, et al.
Publicado: (2023)
por: Bandarkar, Lucas, et al.
Publicado: (2023)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
por: Lin, Yen-Ting, et al.
Publicado: (2025)
por: Lin, Yen-Ting, et al.
Publicado: (2025)
Switching Controller Synthesis for Hybrid Systems Against STL Formulas
por: Su, Han, et al.
Publicado: (2024)
por: Su, Han, et al.
Publicado: (2024)
Additivity of disjoint interval entanglement in quasiparticle excited states
por: Guo, Zhouhao, et al.
Publicado: (2026)
por: Guo, Zhouhao, et al.
Publicado: (2026)
Production and carbon emission reduction decisions of remanufacturing firms with low‐carbon credit financing under uncertain demand
por: Weida Chen, et al.
Publicado: (2024)
por: Weida Chen, et al.
Publicado: (2024)
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
por: Shen, Han, et al.
Publicado: (2024)
por: Shen, Han, et al.
Publicado: (2024)
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
por: Rangarajan, Ajay Mandyam, et al.
Publicado: (2026)
por: Rangarajan, Ajay Mandyam, et al.
Publicado: (2026)
Fabrication and Characterization of Starch/ PVA Blend Films Reinforced With Black Tea for Packaging Applications
por: Vandana Arya, et al.
Publicado: (2026)
por: Vandana Arya, et al.
Publicado: (2026)
On Uniformly Perfect Morse Boundaries
por: Han, Suzhen, et al.
Publicado: (2026)
por: Han, Suzhen, et al.
Publicado: (2026)
WPO: Enhancing RLHF with Weighted Preference Optimization
por: Zhou, Wenxuan, et al.
Publicado: (2024)
por: Zhou, Wenxuan, et al.
Publicado: (2024)
Mind Map of Database
por: Shaw, Karishma
Publicado: (2026)
por: Shaw, Karishma
Publicado: (2026)
Using Agentic AI to Achieve Full CI/CD: A Semantic Reasoning Framework for Microservices Delivery at Scale
por: Karishma Verma
Publicado: (2026)
por: Karishma Verma
Publicado: (2026)
Cardiometabolic Risk Factors in South Asians: An Epidemiological and Anthropological Study in an Urban Populace of Eastern India
por: Yasmin, Karishma
Publicado: (2024)
por: Yasmin, Karishma
Publicado: (2024)
DynaGRAG | Exploring the Topology of Information for Advancing Language Understanding and Generation in Graph Retrieval-Augmented Generation
por: Thakrar, Karishma
Publicado: (2024)
por: Thakrar, Karishma
Publicado: (2024)
Redefining Contributions: Shapley-Driven Federated Learning
por: Tastan, Nurbek, et al.
Publicado: (2024)
por: Tastan, Nurbek, et al.
Publicado: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Dynamic Motion Blending for Versatile Motion Editing
por: Jiang, Nan, et al.
Publicado: (2025)
por: Jiang, Nan, et al.
Publicado: (2025)
Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic
por: Rao, Abinav, et al.
Publicado: (2026)
por: Rao, Abinav, et al.
Publicado: (2026)
Graph Hopfield Networks: Energy-Based Node Classification with Associative Memory
por: Rao, Abinav, et al.
Publicado: (2026)
por: Rao, Abinav, et al.
Publicado: (2026)
Ejemplares similares
-
Reinforcement Learning from User Feedback
por: Han, Eric, et al.
Publicado: (2025) -
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
por: Yu, Zishun, et al.
Publicado: (2025) -
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
por: Qin, Chengwei, et al.
Publicado: (2025) -
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
por: He, Yun, et al.
Publicado: (2024) -
Generalized Parallel Scaling with Interdependent Generations
por: Dong, Harry, et al.
Publicado: (2025)