Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heye, David, Kindermann, Karl, Decker, Robin, Lohmöller, Johannes, Belova, Anastasiia, Geisler, Sandra, Wehrle, Klaus, Pennekamp, Jan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918386015928320
author Heye, David
Kindermann, Karl
Decker, Robin
Lohmöller, Johannes
Belova, Anastasiia
Geisler, Sandra
Wehrle, Klaus
Pennekamp, Jan
author_facet Heye, David
Kindermann, Karl
Decker, Robin
Lohmöller, Johannes
Belova, Anastasiia
Geisler, Sandra
Wehrle, Klaus
Pennekamp, Jan
contents Artifact Evaluation (AE) is essential for ensuring the transparency and reliability of research, closing the gap between exploratory work and real-world deployment is particularly important in cybersecurity, particularly in IoT and CPSs, where large-scale, heterogeneous, and privacy-sensitive data meet safety-critical actuation. Yet, manual reproducibility checks are time-consuming and do not scale with growing submission volumes. In this work, we demonstrate that Large Language Models (LLMs) can provide powerful support for AE tasks: (i) text-based reproducibility rating, (ii) autonomous sandboxed execution environment preparation, and (iii) assessment of methodological pitfalls. Our reproducibility-assessment toolkit yields an accuracy of over 72% and autonomously sets up execution environments for 28% of runnable cybersecurity artifacts. Our automated pitfall assessment detects seven prevalent pitfalls with high accuracy ($F_1$ > 92%). Hence, the toolkit significantly reduces reviewer effort and, when integrated into established AE processes, could incentivize authors to submit higher-quality and more reproducible artifacts. IoT, CPS, and cybersecurity conferences and workshops may integrate the toolkit into their peer-review processes to support reviewers' decisions on awarding artifact badges, improving the overall sustainability of the process.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06862
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
Heye, David
Kindermann, Karl
Decker, Robin
Lohmöller, Johannes
Belova, Anastasiia
Geisler, Sandra
Wehrle, Klaus
Pennekamp, Jan
Cryptography and Security
Artificial Intelligence
Computation and Language
Artifact Evaluation (AE) is essential for ensuring the transparency and reliability of research, closing the gap between exploratory work and real-world deployment is particularly important in cybersecurity, particularly in IoT and CPSs, where large-scale, heterogeneous, and privacy-sensitive data meet safety-critical actuation. Yet, manual reproducibility checks are time-consuming and do not scale with growing submission volumes. In this work, we demonstrate that Large Language Models (LLMs) can provide powerful support for AE tasks: (i) text-based reproducibility rating, (ii) autonomous sandboxed execution environment preparation, and (iii) assessment of methodological pitfalls. Our reproducibility-assessment toolkit yields an accuracy of over 72% and autonomously sets up execution environments for 28% of runnable cybersecurity artifacts. Our automated pitfall assessment detects seven prevalent pitfalls with high accuracy ($F_1$ > 92%). Hence, the toolkit significantly reduces reviewer effort and, when integrated into established AE processes, could incentivize authors to submit higher-quality and more reproducible artifacts. IoT, CPS, and cybersecurity conferences and workshops may integrate the toolkit into their peer-review processes to support reviewers' decisions on awarding artifact badges, improving the overall sustainability of the process.
title Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.06862