S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ma, Ruotian, Wang, Peisong, Liu, Cheng, Liu, Xingyan, Chen, Jiaqi, Zhang, Bang, Zhou, Xin, Du, Nan, Li, Jia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
por: Chen, Jiaqi, et al.
Publicado: (2025)
por: Chen, Jiaqi, et al.
Publicado: (2025)
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
por: Sui, Yuan, et al.
Publicado: (2026)
por: Sui, Yuan, et al.
Publicado: (2026)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
por: Xu, Tianyang, et al.
Publicado: (2024)
por: Xu, Tianyang, et al.
Publicado: (2024)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
por: Wang, Peisong, et al.
Publicado: (2025)
por: Wang, Peisong, et al.
Publicado: (2025)
Self-correction is Not An Innate Capability in Language Models
por: Liu, Guangliang, et al.
Publicado: (2024)
por: Liu, Guangliang, et al.
Publicado: (2024)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
por: Li, Junsong, et al.
Publicado: (2025)
por: Li, Junsong, et al.
Publicado: (2025)
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
por: Bensal, Shelly, et al.
Publicado: (2025)
por: Bensal, Shelly, et al.
Publicado: (2025)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
por: Zhang, Xiaoying, et al.
Publicado: (2024)
por: Zhang, Xiaoying, et al.
Publicado: (2024)
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
por: Zhang, Bang, et al.
Publicado: (2025)
por: Zhang, Bang, et al.
Publicado: (2025)
Knowledge Graph Reasoning with Self-supervised Reinforcement Learning
por: Ma, Ying, et al.
Publicado: (2024)
por: Ma, Ying, et al.
Publicado: (2024)
SSRL: Self-Search Reinforcement Learning
por: Fan, Yuchen, et al.
Publicado: (2025)
por: Fan, Yuchen, et al.
Publicado: (2025)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
por: Du, Dayou, et al.
Publicado: (2024)
por: Du, Dayou, et al.
Publicado: (2024)
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
por: Tan, Hexiang, et al.
Publicado: (2025)
por: Tan, Hexiang, et al.
Publicado: (2025)
Are Large Language Models Good Prompt Optimizers?
por: Ma, Ruotian, et al.
Publicado: (2024)
por: Ma, Ruotian, et al.
Publicado: (2024)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
por: Huang, Yue, et al.
Publicado: (2025)
por: Huang, Yue, et al.
Publicado: (2025)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
por: Wu, Tong, et al.
Publicado: (2025)
por: Wu, Tong, et al.
Publicado: (2025)
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
por: Zhou, Yuxuan, et al.
Publicado: (2025)
por: Zhou, Yuxuan, et al.
Publicado: (2025)
Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
por: Zhang, Haozhen, et al.
Publicado: (2025)
por: Zhang, Haozhen, et al.
Publicado: (2025)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
por: Jiang, Juyong, et al.
Publicado: (2026)
por: Jiang, Juyong, et al.
Publicado: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
por: Yang, Wenkai, et al.
Publicado: (2025)
por: Yang, Wenkai, et al.
Publicado: (2025)
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
por: Song, Huatong, et al.
Publicado: (2025)
por: Song, Huatong, et al.
Publicado: (2025)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
por: Liu, Chi, et al.
Publicado: (2026)
por: Liu, Chi, et al.
Publicado: (2026)
RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
por: Xu, Shicheng, et al.
Publicado: (2025)
por: Xu, Shicheng, et al.
Publicado: (2025)
Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning
por: Ziegenbein, Timon, et al.
Publicado: (2026)
por: Ziegenbein, Timon, et al.
Publicado: (2026)
Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction
por: Feng, Xuan, et al.
Publicado: (2024)
por: Feng, Xuan, et al.
Publicado: (2024)
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
por: Chen, Yongchao, et al.
Publicado: (2025)
por: Chen, Yongchao, et al.
Publicado: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
por: Yang, Xin, et al.
Publicado: (2026)
por: Yang, Xin, et al.
Publicado: (2026)
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
por: Zhou, Ruiyang, et al.
Publicado: (2025)
por: Zhou, Ruiyang, et al.
Publicado: (2025)
Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
por: Xu, Xin, et al.
Publicado: (2025)
por: Xu, Xin, et al.
Publicado: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
por: DeepSeek-AI, et al.
Publicado: (2025)
por: DeepSeek-AI, et al.
Publicado: (2025)
PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning
por: Li, Bingxuan, et al.
Publicado: (2026)
por: Li, Bingxuan, et al.
Publicado: (2026)
Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
por: Kong, Aobo, et al.
Publicado: (2024)
por: Kong, Aobo, et al.
Publicado: (2024)
Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving
por: Zhang, Jin, et al.
Publicado: (2025)
por: Zhang, Jin, et al.
Publicado: (2025)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
por: Zhang, Nonghai, et al.
Publicado: (2026)
por: Zhang, Nonghai, et al.
Publicado: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
por: Zhang, Xiaoying, et al.
Publicado: (2024)
por: Zhang, Xiaoying, et al.
Publicado: (2024)
Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
por: Xu, Jingxin, et al.
Publicado: (2025)
por: Xu, Jingxin, et al.
Publicado: (2025)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
por: Liu, Yilun, et al.
Publicado: (2026)
por: Liu, Yilun, et al.
Publicado: (2026)
Self-Debias: Self-correcting for Debiasing Large Language Models
por: Feng, Xuan, et al.
Publicado: (2026)
por: Feng, Xuan, et al.
Publicado: (2026)
Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
por: Liu, Weize, et al.
Publicado: (2023)
por: Liu, Weize, et al.
Publicado: (2023)
Ejemplares similares
-
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
por: Chen, Jiaqi, et al.
Publicado: (2025) -
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
por: Sui, Yuan, et al.
Publicado: (2026) -
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
por: Xu, Tianyang, et al.
Publicado: (2024) -
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
por: Wang, Peisong, et al.
Publicado: (2025) -
Self-correction is Not An Innate Capability in Language Models
por: Liu, Guangliang, et al.
Publicado: (2024)