AutoDFBench 1.0: A Benchmarking Framework for Digital Forensic Tool Testing and Generated Code Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Wickramasekara, Akila, Mihiranga, Tharusha, Withanage, Aruna, Weerasinghe, Buddhima, Breitinger, Frank, Sheppard, John, Scanlon, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
by: Wickramasekara, Akila, et al.
Published: (2024)
by: Wickramasekara, Akila, et al.
Published: (2024)
Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
by: Michelet, Gaëtan, et al.
Published: (2025)
by: Michelet, Gaëtan, et al.
Published: (2025)
DFRWS EU 10-Year Review and Future Directions in Digital Forensic Research
by: Breitinger, Frank, et al.
Published: (2023)
by: Breitinger, Frank, et al.
Published: (2023)
An Empirical Study of Code Obfuscation Practices in the Google Play Store
by: Niroshan, Akila, et al.
Published: (2025)
by: Niroshan, Akila, et al.
Published: (2025)
Towards a standardized methodology and dataset for evaluating LLM-based digital forensic timeline analysis
by: Studiawan, Hudan, et al.
Published: (2025)
by: Studiawan, Hudan, et al.
Published: (2025)
Defining Atomicity (and Integrity) for Snapshots of Storage in Forensic Computing
by: Ottmann, Jenny, et al.
Published: (2025)
by: Ottmann, Jenny, et al.
Published: (2025)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Strategies and Challenges of Timestamp Tampering for Improved Digital Forensic Event Reconstruction (extended version)
by: Vanini, Céline, et al.
Published: (2024)
by: Vanini, Céline, et al.
Published: (2024)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
by: Xing, Hengrui, et al.
Published: (2025)
by: Xing, Hengrui, et al.
Published: (2025)
Many Tools, Few Exploitable Vulnerabilities: A Survey of 246 Static Code Analyzers for Security
by: Hermann, Kevin, et al.
Published: (2026)
by: Hermann, Kevin, et al.
Published: (2026)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
by: Tessa, Melissa, et al.
Published: (2026)
by: Tessa, Melissa, et al.
Published: (2026)
What's Next, Cloud? A Forensic Framework for Analyzing Self-Hosted Cloud Storage Solutions
by: Külper, Michael, et al.
Published: (2025)
by: Külper, Michael, et al.
Published: (2025)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
An Extensive Comparison of Static Application Security Testing Tools
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test
by: Haque, Mohd Ariful, et al.
Published: (2026)
by: Haque, Mohd Ariful, et al.
Published: (2026)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
Uncovering EDK2 Firmware Flaws: Insights from Code Audit Tools
by: Farahani, Mahsa, et al.
Published: (2024)
by: Farahani, Mahsa, et al.
Published: (2024)
Comparing Effectiveness and Efficiency of Interactive Application Security Testing (IAST) and Runtime Application Self-Protection (RASP) Tools in a Large Java-based System
by: Seth, Aishwarya, et al.
Published: (2023)
by: Seth, Aishwarya, et al.
Published: (2023)
CAShift: Benchmarking Log-Based Cloud Attack Detection under Normality Shift
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
AutoFirm: Automatically Identifying Reused Libraries inside IoT Firmware at Large-Scale
by: Chen, YongLe, et al.
Published: (2024)
by: Chen, YongLe, et al.
Published: (2024)
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models
by: Bruni, Marc, et al.
Published: (2025)
by: Bruni, Marc, et al.
Published: (2025)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
by: Pathak, Abhijeet, et al.
Published: (2025)
by: Pathak, Abhijeet, et al.
Published: (2025)
WildCode: An Empirical Analysis of Code Generated by ChatGPT
by: Khanmohammadi, Kobra, et al.
Published: (2025)
by: Khanmohammadi, Kobra, et al.
Published: (2025)
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
How to Compare the Security of Code Written by Humans to LLM-generated Code
by: Balebako, Rebecca, et al.
Published: (2026)
by: Balebako, Rebecca, et al.
Published: (2026)
SecCodePRM: A Process Reward Model for Code Security
by: Yu, Weichen, et al.
Published: (2026)
by: Yu, Weichen, et al.
Published: (2026)
Evaluating Tool Cloning in Agentic-AI Ecosystems
by: Kim, Taein, et al.
Published: (2026)
by: Kim, Taein, et al.
Published: (2026)
Security Testing of RESTful APIs With Test Case Mutation
by: Salva, Sebastien, et al.
Published: (2024)
by: Salva, Sebastien, et al.
Published: (2024)
Unsupervised Binary Code Translation with Application to Code Similarity Detection and Vulnerability Discovery
by: Ahmad, Iftakhar, et al.
Published: (2024)
by: Ahmad, Iftakhar, et al.
Published: (2024)
Auditing MCP Servers for Over-Privileged Tool Capabilities
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
Are AI-assisted Development Tools Immune to Prompt Injection?
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
SBOM Generation Tools in the Python Ecosystem: an In-Detail Analysis
by: Cofano, Serena, et al.
Published: (2024)
by: Cofano, Serena, et al.
Published: (2024)
Similar Items
-
Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
by: Wickramasekara, Akila, et al.
Published: (2024) -
Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
by: Michelet, Gaëtan, et al.
Published: (2025) -
DFRWS EU 10-Year Review and Future Directions in Digital Forensic Research
by: Breitinger, Frank, et al.
Published: (2023) -
An Empirical Study of Code Obfuscation Practices in the Google Play Store
by: Niroshan, Akila, et al.
Published: (2025) -
Towards a standardized methodology and dataset for evaluating LLM-based digital forensic timeline analysis
by: Studiawan, Hudan, et al.
Published: (2025)