Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Sirui, Xu, Lei, Zhao, Yuying, Chen, Yutian, Wang, Yu, Zhu, Beier, Zhang, Hanwang, Zhao, Shengjie, Lu, Chaochao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Synthesis by Design: Controlled Data Generation via Structural Guidance
by: Xu, Lei, et al.
Published: (2025)
by: Xu, Lei, et al.
Published: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
From Imitation to Introspection: Probing Self-Consciousness in Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
by: Zhang, Yize, et al.
Published: (2025)
by: Zhang, Yize, et al.
Published: (2025)
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability
by: Han, Yujin, et al.
Published: (2024)
by: Han, Yujin, et al.
Published: (2024)
Causal Evaluation of Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
by: Zhao, Tianzhe, et al.
Published: (2026)
by: Zhao, Tianzhe, et al.
Published: (2026)
KALE: Enhancing Knowledge Manipulation in Large Language Models via Knowledge-aware Learning
by: Lv, Qitan, et al.
Published: (2026)
by: Lv, Qitan, et al.
Published: (2026)
Enhancing LLM Reasoning with Reward-guided Tree Search
by: Jiang, Jinhao, et al.
Published: (2024)
by: Jiang, Jinhao, et al.
Published: (2024)
Chain-of-Knowledge: Integrating Knowledge Reasoning into Large Language Models by Learning from Knowledge Graphs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
by: Zhou, Jiang, et al.
Published: (2026)
by: Zhou, Jiang, et al.
Published: (2026)
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
by: Chen, Yanyu, et al.
Published: (2026)
by: Chen, Yanyu, et al.
Published: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
by: Kansal, Yuval, et al.
Published: (2026)
by: Kansal, Yuval, et al.
Published: (2026)
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
by: Fan, Lishui, et al.
Published: (2025)
by: Fan, Lishui, et al.
Published: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
by: Xu, Hongling, et al.
Published: (2025)
by: Xu, Hongling, et al.
Published: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
by: Wang, Sirui, et al.
Published: (2025)
by: Wang, Sirui, et al.
Published: (2025)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
by: Zheng, Hang, et al.
Published: (2025)
by: Zheng, Hang, et al.
Published: (2025)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
by: Zhao, Qingfei, et al.
Published: (2025)
by: Zhao, Qingfei, et al.
Published: (2025)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
by: Zhang, Xiaoying, et al.
Published: (2025)
by: Zhang, Xiaoying, et al.
Published: (2025)
Reasoning Factual Knowledge in Structured Data with Large Language Models
by: Huang, Sirui, et al.
Published: (2024)
by: Huang, Sirui, et al.
Published: (2024)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
by: Xie, Tianbao, et al.
Published: (2023)
by: Xie, Tianbao, et al.
Published: (2023)
Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
by: Sun, Zexu, et al.
Published: (2025)
by: Sun, Zexu, et al.
Published: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Metacognitive Prompting Improves Understanding in Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
by: Shan, Liang, et al.
Published: (2025)
by: Shan, Liang, et al.
Published: (2025)
Reinforced Efficient Reasoning via Semantically Diverse Exploration
by: Zhao, Ziqi, et al.
Published: (2026)
by: Zhao, Ziqi, et al.
Published: (2026)
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
by: Yu, Fangxu, et al.
Published: (2025)
by: Yu, Fangxu, et al.
Published: (2025)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
by: Peng, Hao, et al.
Published: (2025)
by: Peng, Hao, et al.
Published: (2025)
Similar Items
-
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025) -
Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
by: Chen, Sirui, et al.
Published: (2025) -
Synthesis by Design: Controlled Data Generation via Structural Guidance
by: Xu, Lei, et al.
Published: (2025) -
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026) -
From Imitation to Introspection: Probing Self-Consciousness in Language Models
by: Chen, Sirui, et al.
Published: (2024)