CARL: Criticality-Aware Agentic Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Leyang, Zhang, Yang, Ling, Chun Kai, Zhao, Xiaoyan, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913111943938048
author Shen, Leyang
Zhang, Yang
Ling, Chun Kai
Zhao, Xiaoyan
Chua, Tat-Seng
author_facet Shen, Leyang
Zhang, Yang
Ling, Chun Kai
Zhao, Xiaoyan
Chua, Tat-Seng
contents Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction. However, in such multi-step settings, the conventional group-level policy optimization algorithm becomes suboptimal because of its underlying assumption that each step holds equal contribution, which deviates significantly from reality. Our analysis reveals that only the action choices on a small fraction of states are critical in determining the final outcome. Building on this insight, we propose CARL, a criticality-aware reinforcement learning algorithm tailored for long-horizon agentic reasoning. CARL leverages entropy as a heuristic proxy for state criticality and achieves focused training by assigning rewards to actions taken from high-criticality states while excluding actions taken from low-criticality states from model updates, avoiding noisy credit assignment and redundant computation. Extensive experiments demonstrate that CARL achieves both stronger performance and higher efficiency across diverse evaluation settings. The source code will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CARL: Criticality-Aware Agentic Reinforcement Learning
Shen, Leyang
Zhang, Yang
Ling, Chun Kai
Zhao, Xiaoyan
Chua, Tat-Seng
Machine Learning
Artificial Intelligence
Computation and Language
Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction. However, in such multi-step settings, the conventional group-level policy optimization algorithm becomes suboptimal because of its underlying assumption that each step holds equal contribution, which deviates significantly from reality. Our analysis reveals that only the action choices on a small fraction of states are critical in determining the final outcome. Building on this insight, we propose CARL, a criticality-aware reinforcement learning algorithm tailored for long-horizon agentic reasoning. CARL leverages entropy as a heuristic proxy for state criticality and achieves focused training by assigning rewards to actions taken from high-criticality states while excluding actions taken from low-criticality states from model updates, avoiding noisy credit assignment and redundant computation. Extensive experiments demonstrate that CARL achieves both stronger performance and higher efficiency across diverse evaluation settings. The source code will be publicly available.
title CARL: Criticality-Aware Agentic Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.04949