Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jiaye, Li, Mengdi, Zhao, Xufeng, Lu, Wenhao, Zhao, Peilin, Wermter, Stefan, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal State Distillation for Explainable Reinforcement Learning
by: Lu, Wenhao, et al.
Published: (2023)
by: Lu, Wenhao, et al.
Published: (2023)
Large Language Models for Orchestrating Bimanual Robots
by: Chu, Kun, et al.
Published: (2024)
by: Chu, Kun, et al.
Published: (2024)
Mental Modeling of Reinforcement Learning Agents by Language Models
by: Lu, Wenhao, et al.
Published: (2024)
by: Lu, Wenhao, et al.
Published: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
by: Lee, Harrison, et al.
Published: (2023)
by: Lee, Harrison, et al.
Published: (2023)
Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
by: Zhao, Xufeng, et al.
Published: (2023)
by: Zhao, Xufeng, et al.
Published: (2023)
Agentic Skill Discovery
by: Zhao, Xufeng, et al.
Published: (2024)
by: Zhao, Xufeng, et al.
Published: (2024)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
by: Yang, Qing, et al.
Published: (2025)
by: Yang, Qing, et al.
Published: (2025)
LLM+MAP: Bimanual Robot Task Planning using Large Language Models and Planning Domain Definition Language
by: Chu, Kun, et al.
Published: (2025)
by: Chu, Kun, et al.
Published: (2025)
Details Make a Difference: Object State-Sensitive Neurorobotic Task Planning
by: Sun, Xiaowen, et al.
Published: (2024)
by: Sun, Xiaowen, et al.
Published: (2024)
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
by: Li, Mengdi, et al.
Published: (2025)
by: Li, Mengdi, et al.
Published: (2025)
Curriculum Learning for Safety Alignment
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
by: Beck, Jacob
Published: (2025)
by: Beck, Jacob
Published: (2025)
A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering
by: Zhang, Chenliang, et al.
Published: (2025)
by: Zhang, Chenliang, et al.
Published: (2025)
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
by: Rui, Shaohao, et al.
Published: (2025)
by: Rui, Shaohao, et al.
Published: (2025)
Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical Perspective
by: Zhao, Lei, et al.
Published: (2023)
by: Zhao, Lei, et al.
Published: (2023)
Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
by: Yin, Shouyu, et al.
Published: (2026)
by: Yin, Shouyu, et al.
Published: (2026)
Curriculum Abductive Learning
by: Hu, Wen-Chao, et al.
Published: (2025)
by: Hu, Wen-Chao, et al.
Published: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models through Reinforcement Learning from Ranking Feedback
by: Shi, Derek, et al.
Published: (2025)
by: Shi, Derek, et al.
Published: (2025)
Learning Progress Driven Multi-Agent Curriculum
by: Zhao, Wenshuai, et al.
Published: (2022)
by: Zhao, Wenshuai, et al.
Published: (2022)
Towards User-level Private Reinforcement Learning with Human Feedback
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Grounded Curriculum Learning
by: Wang, Linji, et al.
Published: (2024)
by: Wang, Linji, et al.
Published: (2024)
Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
by: Liu, Boyin, et al.
Published: (2025)
by: Liu, Boyin, et al.
Published: (2025)
Probabilistic Curriculum Learning for Goal-Based Reinforcement Learning
by: Salt, Llewyn, et al.
Published: (2025)
by: Salt, Llewyn, et al.
Published: (2025)
AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations
by: Lindström, Adam Dahlgren, et al.
Published: (2024)
by: Lindström, Adam Dahlgren, et al.
Published: (2024)
Learning Versatile Skills with Curriculum Masking
by: Tang, Yao, et al.
Published: (2024)
by: Tang, Yao, et al.
Published: (2024)
Adversarial Curriculum Graph Contrastive Learning with Pair-wise Augmentation
by: Zhao, Xinjian, et al.
Published: (2024)
by: Zhao, Xinjian, et al.
Published: (2024)
AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Human-Aware Robot Navigation via Reinforcement Learning with Hindsight Experience Replay and Curriculum Learning
by: Li, Keyu, et al.
Published: (2021)
by: Li, Keyu, et al.
Published: (2021)
TACLer: Tailored Curriculum Reinforcement Learning for Efficient Reasoning
by: Lai, Huiyuan, et al.
Published: (2026)
by: Lai, Huiyuan, et al.
Published: (2026)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
by: Tzannetos, Georgios, et al.
Published: (2024)
by: Tzannetos, Georgios, et al.
Published: (2024)
Balancing Multiple Objectives in Urban Traffic Control with Reinforcement Learning from AI Feedback
by: Zhao, Chenyang, et al.
Published: (2026)
by: Zhao, Chenyang, et al.
Published: (2026)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
by: Yuan, Ruifeng, et al.
Published: (2025)
by: Yuan, Ruifeng, et al.
Published: (2025)
Exploring Parity Challenges in Reinforcement Learning through Curriculum Learning with Noisy Labels
by: Zhou, Bei, et al.
Published: (2023)
by: Zhou, Bei, et al.
Published: (2023)
Why Does RLAIF Work At All?
by: Young, Robin
Published: (2026)
by: Young, Robin
Published: (2026)
Causally Aligned Curriculum Learning
by: Li, Mingxuan, et al.
Published: (2025)
by: Li, Mingxuan, et al.
Published: (2025)
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
Similar Items
-
Causal State Distillation for Explainable Reinforcement Learning
by: Lu, Wenhao, et al.
Published: (2023) -
Large Language Models for Orchestrating Bimanual Robots
by: Chu, Kun, et al.
Published: (2024) -
Mental Modeling of Reinforcement Learning Agents by Language Models
by: Lu, Wenhao, et al.
Published: (2024) -
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
by: Lee, Harrison, et al.
Published: (2023) -
Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
by: Zhao, Xufeng, et al.
Published: (2023)