RLVF: Learning from Verbal Feedback without Overgeneralization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Stephan, Moritz, Khazatsky, Alexander, Mitchell, Eric, Chen, Annie S, Hsu, Sheryl, Sharma, Archit, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
par: Hsu, Sheryl, et autres
Publié: (2024)
par: Hsu, Sheryl, et autres
Publié: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
par: Sharma, Archit, et autres
Publié: (2024)
par: Sharma, Archit, et autres
Publié: (2024)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
par: Singh, Anikait, et autres
Publié: (2025)
par: Singh, Anikait, et autres
Publié: (2025)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
par: Rafailov, Rafael, et autres
Publié: (2023)
par: Rafailov, Rafael, et autres
Publié: (2023)
Calibrating Language Models with Adaptive Temperature Scaling
par: Xie, Johnathan, et autres
Publié: (2024)
par: Xie, Johnathan, et autres
Publié: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
par: Xie, Johnathan, et autres
Publié: (2024)
par: Xie, Johnathan, et autres
Publié: (2024)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
par: Luo, Renjie, et autres
Publié: (2025)
par: Luo, Renjie, et autres
Publié: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
par: Hejna, Joey, et autres
Publié: (2023)
par: Hejna, Joey, et autres
Publié: (2023)
Reinforcement Learning via Implicit Imitation Guidance
par: Dong, Perry, et autres
Publié: (2025)
par: Dong, Perry, et autres
Publié: (2025)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
par: Shen, Judy Hanwen, et autres
Publié: (2024)
par: Shen, Judy Hanwen, et autres
Publié: (2024)
Inference and Verbalization Functions During In-Context Learning
par: Tao, Junyi, et autres
Publié: (2024)
par: Tao, Junyi, et autres
Publié: (2024)
CURO: Curriculum Learning for Relative Overgeneralization
par: Shi, Lin, et autres
Publié: (2022)
par: Shi, Lin, et autres
Publié: (2022)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
par: Zaman, Kerem, et autres
Publié: (2025)
par: Zaman, Kerem, et autres
Publié: (2025)
Verbal Werewolf: Engage Users with Verbalized Agentic Werewolf Game Framework
par: Fan, Qihui, et autres
Publié: (2025)
par: Fan, Qihui, et autres
Publié: (2025)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
par: Potamitis, Nearchos, et autres
Publié: (2025)
par: Potamitis, Nearchos, et autres
Publié: (2025)
Clarify: Improving Model Robustness With Natural Language Corrections
par: Lee, Yoonho, et autres
Publié: (2024)
par: Lee, Yoonho, et autres
Publié: (2024)
Curating Demonstrations using Online Experience
par: Chen, Annie S., et autres
Publié: (2025)
par: Chen, Annie S., et autres
Publié: (2025)
Stream of Search (SoS): Learning to Search in Language
par: Gandhi, Kanishk, et autres
Publié: (2024)
par: Gandhi, Kanishk, et autres
Publié: (2024)
Effective Structured Prompting by Meta-Learning and Representative Verbalizer
par: Jiang, Weisen, et autres
Publié: (2023)
par: Jiang, Weisen, et autres
Publié: (2023)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
par: Mark, Max Sobol, et autres
Publié: (2024)
par: Mark, Max Sobol, et autres
Publié: (2024)
Affordance-Guided Reinforcement Learning via Visual Prompting
par: Lee, Olivia Y., et autres
Publié: (2024)
par: Lee, Olivia Y., et autres
Publié: (2024)
Verbal Process Supervision Elicits Better Coding Agents
par: Chen, Hao-Yuan, et autres
Publié: (2025)
par: Chen, Hao-Yuan, et autres
Publié: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
par: Qu, Yuxiao, et autres
Publié: (2025)
par: Qu, Yuxiao, et autres
Publié: (2025)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
par: Gao, Jensen, et autres
Publié: (2024)
par: Gao, Jensen, et autres
Publié: (2024)
How do LLMs Compute Verbal Confidence
par: Kumaran, Dharshan, et autres
Publié: (2026)
par: Kumaran, Dharshan, et autres
Publié: (2026)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
par: Shi, Lucy Xiaoyang, et autres
Publié: (2024)
par: Shi, Lucy Xiaoyang, et autres
Publié: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
par: Rafailov, Rafael, et autres
Publié: (2024)
par: Rafailov, Rafael, et autres
Publié: (2024)
Manual Verbalizer Enrichment for Few-Shot Text Classification
par: Nguyen, Quang Anh, et autres
Publié: (2024)
par: Nguyen, Quang Anh, et autres
Publié: (2024)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
par: Sridhar, Ajay, et autres
Publié: (2025)
par: Sridhar, Ajay, et autres
Publié: (2025)
FASTER: Value-Guided Sampling for Fast RL
par: Dong, Perry, et autres
Publié: (2026)
par: Dong, Perry, et autres
Publié: (2026)
Reinforcement Learning with Backtracking Feedback
par: Sel, Bilgehan, et autres
Publié: (2026)
par: Sel, Bilgehan, et autres
Publié: (2026)
EXPO: Stable Reinforcement Learning with Expressive Policies
par: Dong, Perry, et autres
Publié: (2025)
par: Dong, Perry, et autres
Publié: (2025)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
par: Zhang, Long, et autres
Publié: (2026)
par: Zhang, Long, et autres
Publié: (2026)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
par: Zhang, Shun, et autres
Publié: (2024)
par: Zhang, Shun, et autres
Publié: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
par: Lee, Harrison, et autres
Publié: (2023)
par: Lee, Harrison, et autres
Publié: (2023)
Universal Neural Functionals
par: Zhou, Allan, et autres
Publié: (2024)
par: Zhou, Allan, et autres
Publié: (2024)
Parameter Efficient Reinforcement Learning from Human Feedback
par: Sidahmed, Hakim, et autres
Publié: (2024)
par: Sidahmed, Hakim, et autres
Publié: (2024)
Provable Interactive Learning with Hindsight Instruction Feedback
par: Misra, Dipendra, et autres
Publié: (2024)
par: Misra, Dipendra, et autres
Publié: (2024)
Learning Personalized Agents from Human Feedback
par: Liang, Kaiqu, et autres
Publié: (2026)
par: Liang, Kaiqu, et autres
Publié: (2026)
Distributionally Robust Reinforcement Learning with Human Feedback
par: Mandal, Debmalya, et autres
Publié: (2025)
par: Mandal, Debmalya, et autres
Publié: (2025)
Documents similaires
-
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
par: Hsu, Sheryl, et autres
Publié: (2024) -
A Critical Evaluation of AI Feedback for Aligning Large Language Models
par: Sharma, Archit, et autres
Publié: (2024) -
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
par: Singh, Anikait, et autres
Publié: (2025) -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
par: Rafailov, Rafael, et autres
Publié: (2023) -
Calibrating Language Models with Adaptive Temperature Scaling
par: Xie, Johnathan, et autres
Publié: (2024)