Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaoqun, Yu, Feiyang, Li, Xi, Yan, Guanhua, Yang, Ping, Xi, Zhaohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
CyLens: Towards Reinventing Cyber Threat Intelligence in the Paradigm of Agentic Large Language Models
by: Liu, Xiaoqun, et al.
Published: (2025)
by: Liu, Xiaoqun, et al.
Published: (2025)
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
All Your Knowledge Belongs to Us: Stealing Knowledge Graphs via Reasoning APIs
by: Xi, Zhaohan
Published: (2025)
by: Xi, Zhaohan
Published: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
Autonomous Threat Hunting: A Future Paradigm for AI-Driven Threat Intelligence
by: Sindiramutty, Siva Raja
Published: (2023)
by: Sindiramutty, Siva Raja
Published: (2023)
Bandit on the Hunt: Dynamic Crawling for Cyber Threat Intelligence
by: Kuehn, Philipp, et al.
Published: (2025)
by: Kuehn, Philipp, et al.
Published: (2025)
Cyber Threat Hunting: Non-Parametric Mining of Attack Patterns from Cyber Threat Intelligence for Precise Threats Attribution
by: Kanwal, Rimsha, et al.
Published: (2025)
by: Kanwal, Rimsha, et al.
Published: (2025)
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
by: Abuadbba, Alsharif, et al.
Published: (2025)
by: Abuadbba, Alsharif, et al.
Published: (2025)
Can LLMs Threaten Human Survival? Benchmarking Potential Existential Threats from LLMs via Prefix Completion
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Securing Federated Learning against Backdoor Threats with Foundation Model Integration
by: Bi, Xiaohuan, et al.
Published: (2024)
by: Bi, Xiaohuan, et al.
Published: (2024)
Enhancing Cyber Threat Hunting -- A Visual Approach with the Forensic Visualization Toolkit
by: Najar, Jihane, et al.
Published: (2025)
by: Najar, Jihane, et al.
Published: (2025)
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
by: Desai, Pratyush, et al.
Published: (2026)
by: Desai, Pratyush, et al.
Published: (2026)
ActMiner: Applying Causality Tracking and Increment Aligning for Graph-based Cyber Threat Hunting
by: Ma, Mingjun, et al.
Published: (2025)
by: Ma, Mingjun, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks
by: Mu, Yanming, et al.
Published: (2026)
by: Mu, Yanming, et al.
Published: (2026)
APT-CGLP: Advanced Persistent Threat Hunting via Contrastive Graph-Language Pre-Training
by: Qiu, Xuebo, et al.
Published: (2025)
by: Qiu, Xuebo, et al.
Published: (2025)
Rethinking Security in Semantic Communication: Latent Manipulation as a New Threat
by: Xi, Zhiyuan, et al.
Published: (2025)
by: Xi, Zhiyuan, et al.
Published: (2025)
Using LLMs to Automate Threat Intelligence Analysis Workflows in Security Operation Centers
by: Tseng, PeiYu, et al.
Published: (2024)
by: Tseng, PeiYu, et al.
Published: (2024)
Blue Teaming Function-Calling Agents
by: Dolcetti, Greta, et al.
Published: (2026)
by: Dolcetti, Greta, et al.
Published: (2026)
Cybersecurity Threat Hunting and Vulnerability Analysis Using a Neo4j Graph Database of Open Source Intelligence
by: Pelofske, Elijah, et al.
Published: (2023)
by: Pelofske, Elijah, et al.
Published: (2023)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
Technique Inference Engine: A Recommender Model to Support Cyber Threat Hunting
by: Turner, Matthew J., et al.
Published: (2025)
by: Turner, Matthew J., et al.
Published: (2025)
Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage
by: Sahay, Rishikesh, et al.
Published: (2026)
by: Sahay, Rishikesh, et al.
Published: (2026)
Unsupervised Threat Hunting using Continuous Bag-of-Terms-and-Time (CBoTT)
by: Kayhan, Varol, et al.
Published: (2024)
by: Kayhan, Varol, et al.
Published: (2024)
CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
by: Fumero, Stefano, et al.
Published: (2025)
by: Fumero, Stefano, et al.
Published: (2025)
Towards Explainable and Lightweight AI for Real-Time Cyber Threat Hunting in Edge Networks
by: Rahmati, Milad
Published: (2025)
by: Rahmati, Milad
Published: (2025)
ThreatPilot: Attack-Driven Threat Intelligence Extraction
by: Xu, Ming, et al.
Published: (2024)
by: Xu, Ming, et al.
Published: (2024)
SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
by: Ji, Hangyuan, et al.
Published: (2024)
by: Ji, Hangyuan, et al.
Published: (2024)
Securing the Future: Proactive Threat Hunting for Sustainable IoT Ecosystems
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
An Interview Study on Third-Party Cyber Threat Hunting Processes in the U.S. Department of Homeland Security
by: Maxam III, William P., et al.
Published: (2024)
by: Maxam III, William P., et al.
Published: (2024)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
How to Design a Blue Team Scenario for Beginners on the Example of Brute-Force Attacks on Authentications
by: Eipper, Andreas, et al.
Published: (2024)
by: Eipper, Andreas, et al.
Published: (2024)
Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
by: Suo, Pan, et al.
Published: (2025)
by: Suo, Pan, et al.
Published: (2025)
Similar Items
-
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025) -
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025) -
CyLens: Towards Reinventing Cyber Threat Intelligence in the Paradigm of Agentic Large Language Models
by: Liu, Xiaoqun, et al.
Published: (2025) -
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024) -
All Your Knowledge Belongs to Us: Stealing Knowledge Graphs via Reasoning APIs
by: Xi, Zhaohan
Published: (2025)