RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Peisong, Ma, Ruotian, Zhang, Bang, Chen, Xingyu, He, Zhiwei, Luo, Kang, Lv, Qingsong, Jiang, Qingxuan, Xie, Zheng, Wang, Shanyi, Li, Yuan, Ye, Fanghua, Li, Jian, Yang, Yifan, Tu, Zhaopeng, Li, Xiaolong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
von: Zhang, Bang, et al.
Veröffentlicht: (2025)
von: Zhang, Bang, et al.
Veröffentlicht: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
von: Wang, Yue, et al.
Veröffentlicht: (2025)
von: Wang, Yue, et al.
Veröffentlicht: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2023)
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
von: Ma, Xinbei, et al.
Veröffentlicht: (2025)
von: Ma, Xinbei, et al.
Veröffentlicht: (2025)
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
Emotional Support with LLM-based Empathetic Dialogue Generation
von: Wang, Shiquan, et al.
Veröffentlicht: (2025)
von: Wang, Shiquan, et al.
Veröffentlicht: (2025)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
von: He, Zhiwei, et al.
Veröffentlicht: (2024)
von: He, Zhiwei, et al.
Veröffentlicht: (2024)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
von: Guo, Xu, et al.
Veröffentlicht: (2025)
von: Guo, Xu, et al.
Veröffentlicht: (2025)
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
von: Wu, Fang, et al.
Veröffentlicht: (2025)
von: Wu, Fang, et al.
Veröffentlicht: (2025)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
von: Huang, Guanhua, et al.
Veröffentlicht: (2025)
von: Huang, Guanhua, et al.
Veröffentlicht: (2025)
Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
von: Wang, Chengbing, et al.
Veröffentlicht: (2026)
von: Wang, Chengbing, et al.
Veröffentlicht: (2026)
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
von: Jiao, Siwen, et al.
Veröffentlicht: (2026)
von: Jiao, Siwen, et al.
Veröffentlicht: (2026)
CTSM: Combining Trait and State Emotions for Empathetic Response Model
von: Yufeng, Wang, et al.
Veröffentlicht: (2024)
von: Yufeng, Wang, et al.
Veröffentlicht: (2024)
CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming
von: Wang, Peisong, et al.
Veröffentlicht: (2026)
von: Wang, Peisong, et al.
Veröffentlicht: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
APTNESS: Incorporating Appraisal Theory and Emotion Support Strategies for Empathetic Response Generation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
von: Wang, Li, et al.
Veröffentlicht: (2026)
von: Wang, Li, et al.
Veröffentlicht: (2026)
Empathy Level Alignment via Reinforcement Learning for Empathetic Response Generation
von: Ma, Hui, et al.
Veröffentlicht: (2024)
von: Ma, Hui, et al.
Veröffentlicht: (2024)
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
Exploiting Emotion-Semantic Correlations for Empathetic Response Generation
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
von: Ding, Yuyang, et al.
Veröffentlicht: (2025)
von: Ding, Yuyang, et al.
Veröffentlicht: (2025)
CARE: Causality Reasoning for Empathetic Responses by Conditional Graph Generation
von: Wang, Jiashuo, et al.
Veröffentlicht: (2022)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2022)
Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents
von: K, Deeraj S, et al.
Veröffentlicht: (2026)
von: K, Deeraj S, et al.
Veröffentlicht: (2026)
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
Towards Cost-Effective Reward Guided Text Generation
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
von: Wen, Xumeng, et al.
Veröffentlicht: (2025)
von: Wen, Xumeng, et al.
Veröffentlicht: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
von: Zhang, Bang, et al.
Veröffentlicht: (2025) -
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
von: Liu, Cheng, et al.
Veröffentlicht: (2025) -
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025) -
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025) -
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)