Self-Rewarding Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Weizhe, Pang, Richard Yuanzhe, Cho, Kyunghyun, Li, Xian, Sukhbaatar, Sainbayar, Xu, Jing, Weston, Jason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
System-Level Natural Language Feedback
von: Yuan, Weizhe, et al.
Veröffentlicht: (2023)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2023)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
von: Xu, Jing, et al.
Veröffentlicht: (2023)
von: Xu, Jing, et al.
Veröffentlicht: (2023)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Following Length Constraints in Instructions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Self-Taught Evaluators
von: Wang, Tianlu, et al.
Veröffentlicht: (2024)
von: Wang, Tianlu, et al.
Veröffentlicht: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
An Overview of Large Language Models for Statisticians
von: Ji, Wenlong, et al.
Veröffentlicht: (2025)
von: Ji, Wenlong, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Multi-Token Attention
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
Leveraging Implicit Feedback from Deployment Data in Dialogue
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2023)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2023)
SPICE: Self-Play In Corpus Environments Improves Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
Training Large Language Models to Reason in a Continuous Latent Space
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
Self-Generated Critiques Boost Reward Modeling for Language Models
von: Yu, Yue, et al.
Veröffentlicht: (2024)
von: Yu, Yue, et al.
Veröffentlicht: (2024)
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
Distilling System 2 into System 1
von: Yu, Ping, et al.
Veröffentlicht: (2024)
von: Yu, Ping, et al.
Veröffentlicht: (2024)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
Language Models as Causal Effect Generators
von: Bynum, Lucius E. J., et al.
Veröffentlicht: (2024)
von: Bynum, Lucius E. J., et al.
Veröffentlicht: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Process-based Self-Rewarding Language Models
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Diverse Preference Optimization
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Training Language Models with Language Feedback at Scale
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
The Era of Real-World Human Interaction: RL from User Conversations
von: Jin, Chuanyang, et al.
Veröffentlicht: (2025)
von: Jin, Chuanyang, et al.
Veröffentlicht: (2025)
Language Imbalance Driven Rewarding for Multilingual Self-improving
von: Yang, Wen, et al.
Veröffentlicht: (2024)
von: Yang, Wen, et al.
Veröffentlicht: (2024)
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
Better Alignment with Instruction Back-and-Forth Translation
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024) -
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024) -
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025) -
System-Level Natural Language Feedback
von: Yuan, Weizhe, et al.
Veröffentlicht: (2023) -
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)