Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911127481352192 |
|---|---|
| author | Zhu, Yongfu Sun, Lin Zhao, Guangxiang Lin, Weihong Zhang, Xiangzheng |
| author_facet | Zhu, Yongfu Sun, Lin Zhao, Guangxiang Lin, Weihong Zhang, Xiangzheng |
| contents | In this work, we introduce Entropy Area Score (EAS), a simple yet effective metric to quantify uncertainty in the answer generation process of reasoning large language models (LLMs). EAS requires neither external models nor repeated sampling, it integrates token-level predictive entropy from the model itself to capture the evolution of uncertainty during generation. Empirical results show that EAS is strongly correlated with answer entropy across models and datasets. In training data selection, EAS identifies high-potential samples and consistently outperforms Pass Rate filtering under equal sample budgets, improving student model accuracy on math benchmarks. EAS is both efficient and interpretable, offering a practical tool for uncertainty modeling and data quality assessment in LLM training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_20384 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Zhu, Yongfu Sun, Lin Zhao, Guangxiang Lin, Weihong Zhang, Xiangzheng Artificial Intelligence In this work, we introduce Entropy Area Score (EAS), a simple yet effective metric to quantify uncertainty in the answer generation process of reasoning large language models (LLMs). EAS requires neither external models nor repeated sampling, it integrates token-level predictive entropy from the model itself to capture the evolution of uncertainty during generation. Empirical results show that EAS is strongly correlated with answer entropy across models and datasets. In training data selection, EAS identifies high-potential samples and consistently outperforms Pass Rate filtering under equal sample budgets, improving student model accuracy on math benchmarks. EAS is both efficient and interpretable, offering a practical tool for uncertainty modeling and data quality assessment in LLM training. |
| title | Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2508.20384 |