Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ziyao, Sun, Guoheng, He, Yexiao, Shen, Zheyu, Tian, Bowei, Li, Ang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908476579512320
author Wang, Ziyao
Sun, Guoheng
He, Yexiao
Shen, Zheyu
Tian, Bowei
Li, Ang
author_facet Wang, Ziyao
Sun, Guoheng
He, Yexiao
Shen, Zheyu
Tian, Bowei
Li, Ang
contents Commercial LLM services often conceal internal reasoning traces while still charging users for every generated token, including those from hidden intermediate steps, raising concerns of token inflation and potential overbilling. This gap underscores the urgent need for reliable token auditing, yet achieving it is far from straightforward: cryptographic verification (e.g., hash-based signature) offers little assurance when providers control the entire execution pipeline, while user-side prediction struggles with the inherent variance of reasoning LLMs, where token usage fluctuates across domains and prompt styles. To bridge this gap, we present PALACE (Predictive Auditing of LLM APIs via Reasoning Token Count Estimation), a user-side framework that estimates hidden reasoning token counts from prompt-answer pairs without access to internal traces. PALACE introduces a GRPO-augmented adaptation module with a lightweight domain router, enabling dynamic calibration across diverse reasoning tasks and mitigating variance in token usage patterns. Experiments on math, coding, medical, and general reasoning benchmarks show that PALACE achieves low relative error and strong prediction accuracy, supporting both fine-grained cost auditing and inflation detection. Taken together, PALACE represents an important first step toward standardized predictive auditing, offering a practical path to greater transparency, accountability, and user trust.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
Wang, Ziyao
Sun, Guoheng
He, Yexiao
Shen, Zheyu
Tian, Bowei
Li, Ang
Machine Learning
Artificial Intelligence
Cryptography and Security
Commercial LLM services often conceal internal reasoning traces while still charging users for every generated token, including those from hidden intermediate steps, raising concerns of token inflation and potential overbilling. This gap underscores the urgent need for reliable token auditing, yet achieving it is far from straightforward: cryptographic verification (e.g., hash-based signature) offers little assurance when providers control the entire execution pipeline, while user-side prediction struggles with the inherent variance of reasoning LLMs, where token usage fluctuates across domains and prompt styles. To bridge this gap, we present PALACE (Predictive Auditing of LLM APIs via Reasoning Token Count Estimation), a user-side framework that estimates hidden reasoning token counts from prompt-answer pairs without access to internal traces. PALACE introduces a GRPO-augmented adaptation module with a lightweight domain router, enabling dynamic calibration across diverse reasoning tasks and mitigating variance in token usage patterns. Experiments on math, coding, medical, and general reasoning benchmarks show that PALACE achieves low relative error and strong prediction accuracy, supporting both fine-grained cost auditing and inflation detection. Taken together, PALACE represents an important first step toward standardized predictive auditing, offering a practical path to greater transparency, accountability, and user trust.
title Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2508.00912