DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response
Fuente:
arXiv
Saved in:
| Main Authors: | Cherif, Bilel, Bisztray, Tamas, Dubniczky, Richard A., Aldahmani, Aaesha, Alshehhi, Saeed, Tihanyi, Norbert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
by: Tihanyi, Norbert, et al.
Published: (2025)
by: Tihanyi, Norbert, et al.
Published: (2025)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Sustaining Cyber Awareness: The Long-Term Impact of Continuous Phishing Training and Emotional Triggers
by: Toth, Rebeka, et al.
Published: (2025)
by: Toth, Rebeka, et al.
Published: (2025)
ForensicsData: A Digital Forensics Dataset for Large Language Models
by: Chakir, Youssef, et al.
Published: (2025)
by: Chakir, Youssef, et al.
Published: (2025)
Software Unclonable Functions for IoT Devices Identification and Security
by: Alshehhi, Saeed
Published: (2025)
by: Alshehhi, Saeed
Published: (2025)
The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
by: Toth, Rebeka, et al.
Published: (2025)
by: Toth, Rebeka, et al.
Published: (2025)
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
by: Tihanyi, Norbert, et al.
Published: (2025)
by: Tihanyi, Norbert, et al.
Published: (2025)
Chances and Challenges of the Model Context Protocol in Digital Forensics and Incident Response
by: Hilgert, Jan-Niclas, et al.
Published: (2025)
by: Hilgert, Jan-Niclas, et al.
Published: (2025)
Real-time Threat Detection Strategies for Resource-constrained Devices
by: Hamidouche, Mounia, et al.
Published: (2024)
by: Hamidouche, Mounia, et al.
Published: (2024)
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era?
by: Bhandarkar, Avanti, et al.
Published: (2024)
by: Bhandarkar, Avanti, et al.
Published: (2024)
IRCopilot: Automated Incident Response with Large Language Models
by: Lin, Xihuan, et al.
Published: (2025)
by: Lin, Xihuan, et al.
Published: (2025)
Digital Forensics in the Age of Large Language Models
by: Yin, Zhipeng, et al.
Published: (2025)
by: Yin, Zhipeng, et al.
Published: (2025)
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
by: Jain, Ridhi, et al.
Published: (2024)
by: Jain, Ridhi, et al.
Published: (2024)
Multi-Agent Collaboration in Incident Response with Large Language Models
by: Liu, Zefang
Published: (2024)
by: Liu, Zefang
Published: (2024)
A Comprehensive Analysis of the Role of Artificial Intelligence and Machine Learning in Modern Digital Forensics and Incident Response
by: Dunsin, Dipo, et al.
Published: (2023)
by: Dunsin, Dipo, et al.
Published: (2023)
Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
by: Wickramasekara, Akila, et al.
Published: (2024)
by: Wickramasekara, Akila, et al.
Published: (2024)
Edge Learning for 6G-enabled Internet of Things: A Comprehensive Survey of Vulnerabilities, Datasets, and Defenses
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
Alignment of Cybersecurity Incident Prioritisation with Incident Response Management Maturity Capabilities
by: Gulay, Abdulaziz, et al.
Published: (2024)
by: Gulay, Abdulaziz, et al.
Published: (2024)
AutoDFBench 1.0: A Benchmarking Framework for Digital Forensic Tool Testing and Generated Code Evaluation
by: Wickramasekara, Akila, et al.
Published: (2025)
by: Wickramasekara, Akila, et al.
Published: (2025)
Digital Forensic Investigation of the ChatGPT Windows Application
by: Kankanamge, Malithi Wanniarachchi, et al.
Published: (2025)
by: Kankanamge, Malithi Wanniarachchi, et al.
Published: (2025)
Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination
by: Hammar, Kim, et al.
Published: (2025)
by: Hammar, Kim, et al.
Published: (2025)
Improving Cybercrime Detection and Digital Forensics Investigations with Artificial Intelligence
by: Sanna, Silvia Lucia, et al.
Published: (2025)
by: Sanna, Silvia Lucia, et al.
Published: (2025)
Conceptualising an Anti-Digital Forensics Kill Chain for Smart Homes
by: Raciti, Mario
Published: (2023)
by: Raciti, Mario
Published: (2023)
Cloud Security Assurance: Strategies for Encryption in Digital Forensic Readiness
by: Alenezi, Ahmed MohanRaj
Published: (2024)
by: Alenezi, Ahmed MohanRaj
Published: (2024)
Employing LLMs for Incident Response Planning and Review
by: Hays, Sam, et al.
Published: (2024)
by: Hays, Sam, et al.
Published: (2024)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
by: Bisztray, Tamas, et al.
Published: (2025)
by: Bisztray, Tamas, et al.
Published: (2025)
In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach
by: Gao, Yiran, et al.
Published: (2026)
by: Gao, Yiran, et al.
Published: (2026)
An Attack-Driven Incident Response and Defense System (ADIRDS)
by: Lai, Anthony Cheuk Tung, et al.
Published: (2025)
by: Lai, Anthony Cheuk Tung, et al.
Published: (2025)
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
by: Zhou, Shide, et al.
Published: (2025)
by: Zhou, Shide, et al.
Published: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
To See or Not to See: A Privacy Threat Model for Digital Forensics in Crime Investigation
by: Raciti, Mario, et al.
Published: (2025)
by: Raciti, Mario, et al.
Published: (2025)
A Smart City Infrastructure Ontology for Threats, Cybercrime, and Digital Forensic Investigation
by: Tok, Yee Ching, et al.
Published: (2024)
by: Tok, Yee Ching, et al.
Published: (2024)
DFRWS EU 10-Year Review and Future Directions in Digital Forensic Research
by: Breitinger, Frank, et al.
Published: (2023)
by: Breitinger, Frank, et al.
Published: (2023)
Towards Lightweight and Privacy-preserving Data Provision in Digital Forensics for Driverless Taxi
by: Gong, Yanwei, et al.
Published: (2024)
by: Gong, Yanwei, et al.
Published: (2024)
Similar Items
-
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
by: Tihanyi, Norbert, et al.
Published: (2025) -
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
by: Dubniczky, Richard A., et al.
Published: (2025) -
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024) -
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025) -
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
by: Ferrag, Mohamed Amine, et al.
Published: (2024)