Gespeichert in:
| Hauptverfasser: | Wang, Tianlu, Kulikov, Ilia, Golovneva, Olga, Yu, Ping, Yuan, Weizhe, Dwivedi-Yu, Jane, Pang, Richard Yuanzhe, Fazel-Zarandi, Maryam, Weston, Jason, Li, Xian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.02666 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Self-Rewarding Language Models
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Distilling System 2 into System 1
von: Yu, Ping, et al.
Veröffentlicht: (2024)
von: Yu, Ping, et al.
Veröffentlicht: (2024)
Multi-Token Attention
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Following Length Constraints in Instructions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
System-Level Natural Language Feedback
von: Yuan, Weizhe, et al.
Veröffentlicht: (2023)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2023)
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2025)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2025)
Efficient Tool Use with Chain-of-Abstraction Reasoning
von: Gao, Silin, et al.
Veröffentlicht: (2024)
von: Gao, Silin, et al.
Veröffentlicht: (2024)
The Era of Real-World Human Interaction: RL from User Conversations
von: Jin, Chuanyang, et al.
Veröffentlicht: (2025)
von: Jin, Chuanyang, et al.
Veröffentlicht: (2025)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Diverse Preference Optimization
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
von: Lupidi, Alisia, et al.
Veröffentlicht: (2024)
von: Lupidi, Alisia, et al.
Veröffentlicht: (2024)
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
Self-Taught Agentic Long Context Understanding
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Leveraging Implicit Feedback from Deployment Data in Dialogue
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2023)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2023)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
TOOLVERIFIER: Generalization to New Tools via Self-Verification
von: Mekala, Dheeraj, et al.
Veröffentlicht: (2024)
von: Mekala, Dheeraj, et al.
Veröffentlicht: (2024)
The Majority is not always right: RL training for solution aggregation
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
von: Yu, Zhaoning, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoning, et al.
Veröffentlicht: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025) -
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024) -
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024) -
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025) -
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)