CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Deep, Badhe, Sanket, Kathrotia, Nehal, Tiwari, Priyanka
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917411456811008
author Shah, Deep
Badhe, Sanket
Kathrotia, Nehal
Tiwari, Priyanka
author_facet Shah, Deep
Badhe, Sanket
Kathrotia, Nehal
Tiwari, Priyanka
contents Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. Existing automatic prompt optimization(APO) frameworks target task accuracy exclusively at the expense of generating long reasoning traces. We propose Cost-Regularized Optimization of Prompts (CROP), an APO method that introduces regularization on response length by generating textual feedback in addition to standard accuracy feedback. This forces the optimization process to produce prompts that elicit concise responses containing only critical information and reasoning. We evaluate our approach on complex reasoning datasets, specifically GSM8K, LogiQA and BIG-Bench Hard. We achieved an 80.6\% reduction in token consumption while maintaining competitive accuracy, seeing only a nominal decline in performance. This presents a pragmatic solution for deploying token-efficient and cost-effective agentic AI systems in production pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14214
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
Shah, Deep
Badhe, Sanket
Kathrotia, Nehal
Tiwari, Priyanka
Computation and Language
Artificial Intelligence
Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. Existing automatic prompt optimization(APO) frameworks target task accuracy exclusively at the expense of generating long reasoning traces. We propose Cost-Regularized Optimization of Prompts (CROP), an APO method that introduces regularization on response length by generating textual feedback in addition to standard accuracy feedback. This forces the optimization process to produce prompts that elicit concise responses containing only critical information and reasoning. We evaluate our approach on complex reasoning datasets, specifically GSM8K, LogiQA and BIG-Bench Hard. We achieved an 80.6\% reduction in token consumption while maintaining competitive accuracy, seeing only a nominal decline in performance. This presents a pragmatic solution for deploying token-efficient and cost-effective agentic AI systems in production pipelines.
title CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.14214