DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hashemi, Masoud, Bamgbose, Oluwanifemi, Madhusudhan, Sathwik Tejaswi, Nair, Jishnu Sethumadhavan, Tiwari, Aman, Yadav, Vikas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917989324947456
author Hashemi, Masoud
Bamgbose, Oluwanifemi
Madhusudhan, Sathwik Tejaswi
Nair, Jishnu Sethumadhavan
Tiwari, Aman
Yadav, Vikas
author_facet Hashemi, Masoud
Bamgbose, Oluwanifemi
Madhusudhan, Sathwik Tejaswi
Nair, Jishnu Sethumadhavan
Tiwari, Aman
Yadav, Vikas
contents Test-time scaling has significantly improved large language model performance, enabling deeper reasoning to solve complex problems. However, this increased reasoning capability also leads to excessive token generation and unnecessary problem-solving attempts. We introduce Dont Reason Bench (DNR Bench), a new benchmark designed to evaluate LLMs ability to robustly understand the tricky reasoning triggers and avoiding unnecessary generation. DNR Bench consists of 150 adversarially designed prompts that are easy for humans to understand and respond to, but surprisingly not for many of the recent prominent LLMs. DNR Bench tests models abilities across different capabilities, such as instruction adherence, hallucination avoidance, redundancy filtering, and unanswerable question recognition. We evaluate reasoning LLMs (RLMs), including DeepSeek-R1, OpenAI O3-mini, Claude-3.7-sonnet and compare them against a powerful non-reasoning model, e.g., GPT-4o. Our experiments reveal that RLMs generate up to 70x more tokens than necessary, often failing at tasks that simpler non-reasoning models handle efficiently with higher accuracy. Our findings underscore the need for more effective training and inference strategies in RLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
Hashemi, Masoud
Bamgbose, Oluwanifemi
Madhusudhan, Sathwik Tejaswi
Nair, Jishnu Sethumadhavan
Tiwari, Aman
Yadav, Vikas
Machine Learning
Test-time scaling has significantly improved large language model performance, enabling deeper reasoning to solve complex problems. However, this increased reasoning capability also leads to excessive token generation and unnecessary problem-solving attempts. We introduce Dont Reason Bench (DNR Bench), a new benchmark designed to evaluate LLMs ability to robustly understand the tricky reasoning triggers and avoiding unnecessary generation. DNR Bench consists of 150 adversarially designed prompts that are easy for humans to understand and respond to, but surprisingly not for many of the recent prominent LLMs. DNR Bench tests models abilities across different capabilities, such as instruction adherence, hallucination avoidance, redundancy filtering, and unanswerable question recognition. We evaluate reasoning LLMs (RLMs), including DeepSeek-R1, OpenAI O3-mini, Claude-3.7-sonnet and compare them against a powerful non-reasoning model, e.g., GPT-4o. Our experiments reveal that RLMs generate up to 70x more tokens than necessary, often failing at tasks that simpler non-reasoning models handle efficiently with higher accuracy. Our findings underscore the need for more effective training and inference strategies in RLMs.
title DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
topic Machine Learning
url https://arxiv.org/abs/2503.15793