LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bekmyradov, Vekil, Pütz, Noah C., Bartz-Beielstein, Thomas
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913036919373824
author Bekmyradov, Vekil
Pütz, Noah C.
Bartz-Beielstein, Thomas
author_facet Bekmyradov, Vekil
Pütz, Noah C.
Bartz-Beielstein, Thomas
contents Large Language Models (LLMs) have achieved impressive results on public benchmarks, often leading to claims of advanced reasoning and understanding. However, recent research in cognitive science reveals that these models sometimes rely on shallow heuristics and memorization, taking shortcuts rather than demonstrating genuine cognitive abilities. This paper investigates LLM behavior in automated test generation for software, contrasting performance on an open-source system (LevelDB) with SAP HANA, one of the most widely deployed commercial database systems worldwide, whose proprietary codebase is guaranteed to be absent from training data. We combine cognitive evaluation principles, drawing on Mitchell's mechanism-focused assessment methodology, with empirical software testing, employing mutation score and iterative compiler-feedback repair loops to assess both accuracy and underlying reasoning strategies. Results show that LLMs excel on familiar, open-source benchmarks but struggle with unseen, complex domains, often prioritizing compilability over semantic effectiveness. These findings provide independent software engineering evidence for the broader claim that current LLMs lack robust reasoning, and highlight the need for evaluation frameworks that penalize trivial shortcuts and reward true generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14437
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
Bekmyradov, Vekil
Pütz, Noah C.
Bartz-Beielstein, Thomas
Software Engineering
Artificial Intelligence
68T07
I.2.7; I.2.1; I.2.5
Large Language Models (LLMs) have achieved impressive results on public benchmarks, often leading to claims of advanced reasoning and understanding. However, recent research in cognitive science reveals that these models sometimes rely on shallow heuristics and memorization, taking shortcuts rather than demonstrating genuine cognitive abilities. This paper investigates LLM behavior in automated test generation for software, contrasting performance on an open-source system (LevelDB) with SAP HANA, one of the most widely deployed commercial database systems worldwide, whose proprietary codebase is guaranteed to be absent from training data. We combine cognitive evaluation principles, drawing on Mitchell's mechanism-focused assessment methodology, with empirical software testing, employing mutation score and iterative compiler-feedback repair loops to assess both accuracy and underlying reasoning strategies. Results show that LLMs excel on familiar, open-source benchmarks but struggle with unseen, complex domains, often prioritizing compilability over semantic effectiveness. These findings provide independent software engineering evidence for the broader claim that current LLMs lack robust reasoning, and highlight the need for evaluation frameworks that penalize trivial shortcuts and reward true generalization.
title LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
topic Software Engineering
Artificial Intelligence
68T07
I.2.7; I.2.1; I.2.5
url https://arxiv.org/abs/2604.14437