OckBench: Measuring the Efficiency of LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Zheng, Kang, Hao, Han, Song, Krishna, Tushar, Zhu, Ligeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914345935437824
author Du, Zheng
Kang, Hao
Han, Song
Krishna, Tushar
Zhu, Ligeng
author_facet Du, Zheng
Kang, Hao
Han, Song
Krishna, Tushar
Zhu, Ligeng
contents Large language models (LLMs) such as GPT-5 and Gemini 3 have pushed the frontier of automated reasoning and code generation. Yet current benchmarks emphasize accuracy and output quality, neglecting a critical dimension: efficiency of token usage. The token efficiency is highly variable in practical. Models solving the same problem with similar accuracy can exhibit up to a \textbf{5.0$\times$} difference in token length, leading to massive gap of model reasoning ability. Such variance exposes significant redundancy, highlighting the critical need for a standardized benchmark to quantify the gap of token efficiency. Thus, we introduce OckBench, the first benchmark that jointly measures accuracy and token efficiency across reasoning and coding tasks. Our evaluation reveals that token efficiency remains largely unoptimized across current models, significantly inflating serving costs and latency. These findings provide a concrete roadmap for the community to optimize the latent reasoning ability, token efficiency. Ultimately, we argue for an evaluation paradigm shift: tokens must not be multiplied beyond necessity. Our benchmarks are available at https://ockbench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05722
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OckBench: Measuring the Efficiency of LLM Reasoning
Du, Zheng
Kang, Hao
Han, Song
Krishna, Tushar
Zhu, Ligeng
Computation and Language
Artificial Intelligence
Large language models (LLMs) such as GPT-5 and Gemini 3 have pushed the frontier of automated reasoning and code generation. Yet current benchmarks emphasize accuracy and output quality, neglecting a critical dimension: efficiency of token usage. The token efficiency is highly variable in practical. Models solving the same problem with similar accuracy can exhibit up to a \textbf{5.0$\times$} difference in token length, leading to massive gap of model reasoning ability. Such variance exposes significant redundancy, highlighting the critical need for a standardized benchmark to quantify the gap of token efficiency. Thus, we introduce OckBench, the first benchmark that jointly measures accuracy and token efficiency across reasoning and coding tasks. Our evaluation reveals that token efficiency remains largely unoptimized across current models, significantly inflating serving costs and latency. These findings provide a concrete roadmap for the community to optimize the latent reasoning ability, token efficiency. Ultimately, we argue for an evaluation paradigm shift: tokens must not be multiplied beyond necessity. Our benchmarks are available at https://ockbench.github.io/.
title OckBench: Measuring the Efficiency of LLM Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.05722