Conformal Constrained Policy Optimization for Cost-Effective LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Wenwen, Jang, Sooyong, Lee, Insup, Bastani, Osbert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914412634308608
author Si, Wenwen
Jang, Sooyong
Lee, Insup
Bastani, Osbert
author_facet Si, Wenwen
Jang, Sooyong
Lee, Insup
Bastani, Osbert
contents While large language models (LLMs) have recently made tremendous progress towards solving challenging AI problems, they have done so at increasingly steep computational and API costs. We propose a novel strategy where we combine multiple LLM models with varying cost/accuracy tradeoffs in an agentic manner, where models and tools are run in sequence as determined by an orchestration model to minimize cost subject to a user-specified level of reliability; this constraint is formalized using conformal prediction to provide guarantees. To solve this problem, we propose Conformal Constrained Policy Optimization (CCPO), a training paradigm that integrates constrained policy optimization with off-policy reinforcement learning and recent advances in online conformal prediction. CCPO jointly optimizes a cost-aware policy (score function) and an adaptive threshold. Across two multi-hop question answering benchmarks, CCPO achieves up to a 30% cost reduction compared to other cost-aware baselines and LLM-guided methods without compromising reliability. Our approach provides a principled and practical framework for deploying LLM agents that are significantly more cost-effective while maintaining reliability.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
Si, Wenwen
Jang, Sooyong
Lee, Insup
Bastani, Osbert
Machine Learning
Artificial Intelligence
While large language models (LLMs) have recently made tremendous progress towards solving challenging AI problems, they have done so at increasingly steep computational and API costs. We propose a novel strategy where we combine multiple LLM models with varying cost/accuracy tradeoffs in an agentic manner, where models and tools are run in sequence as determined by an orchestration model to minimize cost subject to a user-specified level of reliability; this constraint is formalized using conformal prediction to provide guarantees. To solve this problem, we propose Conformal Constrained Policy Optimization (CCPO), a training paradigm that integrates constrained policy optimization with off-policy reinforcement learning and recent advances in online conformal prediction. CCPO jointly optimizes a cost-aware policy (score function) and an adaptive threshold. Across two multi-hop question answering benchmarks, CCPO achieves up to a 30% cost reduction compared to other cost-aware baselines and LLM-guided methods without compromising reliability. Our approach provides a principled and practical framework for deploying LLM agents that are significantly more cost-effective while maintaining reliability.
title Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.11828