CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Keuntae, Jeong, Eunhye, Lee, Sehyeon, Yoon, Seohee, Choi, Yong Suk
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908604226863104
author Kim, Keuntae
Jeong, Eunhye
Lee, Sehyeon
Yoon, Seohee
Choi, Yong Suk
author_facet Kim, Keuntae
Jeong, Eunhye
Lee, Sehyeon
Yoon, Seohee
Choi, Yong Suk
contents Recent advances in enhancing the reasoning ability of large language models (LLMs) have been remarkably successful. LLMs trained with reinforcement learning (RL) for reasoning demonstrate strong performance in challenging tasks such as mathematics and coding, even with relatively small model sizes. However, despite these improvements in task accuracy, the assessment of creativity in LLM generations has been largely overlooked in reasoning tasks, in contrast to writing tasks. The lack of research on creativity assessment in reasoning primarily stems from two challenges: (1) the difficulty of defining the range of creativity, and (2) the necessity of human evaluation in the assessment process. To address these challenges, we propose CLAWS, a method that defines and classifies mathematical solutions into typical, creative, and hallucinated categories without human evaluation, by leveraging attention weights across prompt sections and output. CLAWS outperforms five existing white-box detection methods (Perplexity, Logit Entropy, Window Entropy, Hidden Score, and Attention Score) on five 7-8B math RL models (DeepSeek, Qwen, Mathstral, OpenMath2, and Oreal). We validate CLAWS on 4545 math problems collected from 181 math contests (AJHSME, AMC, AIME).
format Preprint
id arxiv_https___arxiv_org_abs_2510_17921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections
Kim, Keuntae
Jeong, Eunhye
Lee, Sehyeon
Yoon, Seohee
Choi, Yong Suk
Computation and Language
Artificial Intelligence
Recent advances in enhancing the reasoning ability of large language models (LLMs) have been remarkably successful. LLMs trained with reinforcement learning (RL) for reasoning demonstrate strong performance in challenging tasks such as mathematics and coding, even with relatively small model sizes. However, despite these improvements in task accuracy, the assessment of creativity in LLM generations has been largely overlooked in reasoning tasks, in contrast to writing tasks. The lack of research on creativity assessment in reasoning primarily stems from two challenges: (1) the difficulty of defining the range of creativity, and (2) the necessity of human evaluation in the assessment process. To address these challenges, we propose CLAWS, a method that defines and classifies mathematical solutions into typical, creative, and hallucinated categories without human evaluation, by leveraging attention weights across prompt sections and output. CLAWS outperforms five existing white-box detection methods (Perplexity, Logit Entropy, Window Entropy, Hidden Score, and Attention Score) on five 7-8B math RL models (DeepSeek, Qwen, Mathstral, OpenMath2, and Oreal). We validate CLAWS on 4545 math problems collected from 181 math contests (AJHSME, AMC, AIME).
title CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.17921