Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications
Fuente:
arXiv
Saved in:
| Main Author: | Härer, Felix |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024)
by: Muzsai, Lajos, et al.
Published: (2024)
Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
by: Paim, Kayua Oleques, et al.
Published: (2025)
by: Paim, Kayua Oleques, et al.
Published: (2025)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025)
by: Muzsai, Lajos, et al.
Published: (2025)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
by: Schaeffer, Joachim, et al.
Published: (2026)
by: Schaeffer, Joachim, et al.
Published: (2026)
Toward Autonomous and Efficient Cybersecurity: A Multi-Objective AutoML-based Intrusion Detection System
by: Yang, Li, et al.
Published: (2025)
by: Yang, Li, et al.
Published: (2025)
Evaluating the robustness of adversarial defenses in malware detection systems
by: Jafari, Mostafa, et al.
Published: (2025)
by: Jafari, Mostafa, et al.
Published: (2025)
Hacking, The Lazy Way: LLM Augmented Pentesting
by: Goyal, Dhruva, et al.
Published: (2024)
by: Goyal, Dhruva, et al.
Published: (2024)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
On-Premise SLMs vs. Commercial LLMs: Prompt Engineering and Incident Classification in SOCs and CSIRTs
by: Almeida, Gefté, et al.
Published: (2025)
by: Almeida, Gefté, et al.
Published: (2025)
MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation
by: Rocha, Vanderson, et al.
Published: (2025)
by: Rocha, Vanderson, et al.
Published: (2025)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
by: Consoli, Sergio, et al.
Published: (2025)
by: Consoli, Sergio, et al.
Published: (2025)
Towards Autonomous Cybersecurity: An Intelligent AutoML Framework for Autonomous Intrusion Detection
by: Yang, Li, et al.
Published: (2024)
by: Yang, Li, et al.
Published: (2024)
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
by: Charles, Richard M., et al.
Published: (2025)
by: Charles, Richard M., et al.
Published: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Structured Extraction of Vulnerabilities in OpenVAS and Tenable WAS Reports Using LLMs
by: Machado, Beatriz, et al.
Published: (2025)
by: Machado, Beatriz, et al.
Published: (2025)
AnonLFI 2.0: Extensible Architecture for PII Pseudonymization in CSIRTs with OCR and Technical Recognizers
by: Kapelinski, Cristhian, et al.
Published: (2025)
by: Kapelinski, Cristhian, et al.
Published: (2025)
A New Similarity Function for Spectral Clustering with Application to Plant Phenotypic Data
by: Ahuja, Kapil, et al.
Published: (2023)
by: Ahuja, Kapil, et al.
Published: (2023)
Combination of Weak Learners eXplanations to Improve Random Forest eXplicability Robustness
by: Pala, Riccardo, et al.
Published: (2024)
by: Pala, Riccardo, et al.
Published: (2024)
A Multi-Stage Automated Online Network Data Stream Analytics Framework for IIoT Systems
by: Yang, Li, et al.
Published: (2022)
by: Yang, Li, et al.
Published: (2022)
CyberAId: AI-Driven Cybersecurity for Financial Service Providers
by: Fatouros, George, et al.
Published: (2026)
by: Fatouros, George, et al.
Published: (2026)
Blind Gods and Broken Screens: Architecting a Secure, Intent-Centric Mobile Agent Operating System
by: Zou, Zhenhua, et al.
Published: (2026)
by: Zou, Zhenhua, et al.
Published: (2026)
Rethinking Evaluation of Multiple Sclerosis (MS) Lesion Segmentation Models
by: Basit, Abdul, et al.
Published: (2026)
by: Basit, Abdul, et al.
Published: (2026)
MH-1M: A 1.34 Million-Sample Comprehensive Multi-Feature Android Malware Dataset for Machine Learning, Deep Learning, Large Language Models, and Threat Intelligence Research
by: Braganca, Hendrio, et al.
Published: (2025)
by: Braganca, Hendrio, et al.
Published: (2025)
X-NegoBox: An Explainable Privacy-Budget Negotiation Framework for Secure Peer-to-Peer Energy Data Exchange
by: Sengupta, Poushali, et al.
Published: (2026)
by: Sengupta, Poushali, et al.
Published: (2026)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026)
by: Yagoubi, Faouzi El, et al.
Published: (2026)
Pedestrian intention prediction in Adverse Weather Conditions with Spiking Neural Networks and Dynamic Vision Sensors
by: Sakhai, Mustafa, et al.
Published: (2024)
by: Sakhai, Mustafa, et al.
Published: (2024)
A Robust Federated Learning Approach for Combating Attacks Against IoT Systems Under non-IID Challenges
by: Gad, Eyad, et al.
Published: (2025)
by: Gad, Eyad, et al.
Published: (2025)
STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving
by: Hugglestone, James, et al.
Published: (2026)
by: Hugglestone, James, et al.
Published: (2026)
Vision transformer-based multi-camera multi-object tracking framework for dairy cow monitoring
by: Abbas, Kumail, et al.
Published: (2025)
by: Abbas, Kumail, et al.
Published: (2025)
Towards Zero Touch Networks: Cross-Layer Automated Security Solutions for 6G Wireless Networks
by: Yang, Li, et al.
Published: (2025)
by: Yang, Li, et al.
Published: (2025)
Enabling AutoML for Zero-Touch Network Security: Use-Case Driven Analysis
by: Yang, Li, et al.
Published: (2025)
by: Yang, Li, et al.
Published: (2025)
Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning
by: García-Márquez, Mario, et al.
Published: (2026)
by: García-Márquez, Mario, et al.
Published: (2026)
Challenges and Future Directions in Agentic Reverse Engineering Systems
by: Radey, Salem, et al.
Published: (2026)
by: Radey, Salem, et al.
Published: (2026)
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support
by: Svoboda, Igor, et al.
Published: (2024)
by: Svoboda, Igor, et al.
Published: (2024)
Unlocking the Potential of Metaverse in Innovative and Immersive Digital Health
by: Ebrahimzadeh, Fatemeh, et al.
Published: (2024)
by: Ebrahimzadeh, Fatemeh, et al.
Published: (2024)
Attention Please: What Transformer Models Really Learn for Process Prediction
by: Käppel, Martin, et al.
Published: (2024)
by: Käppel, Martin, et al.
Published: (2024)
MalPurifier: Enhancing Android Malware Detection with Adversarial Purification against Evasion Attacks
by: Zhou, Yuyang, et al.
Published: (2023)
by: Zhou, Yuyang, et al.
Published: (2023)
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
by: Ravindran, Santhosh Kumar
Published: (2026)
by: Ravindran, Santhosh Kumar
Published: (2026)
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
by: Yoon, Hyungjun, et al.
Published: (2026)
by: Yoon, Hyungjun, et al.
Published: (2026)
Similar Items
-
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024) -
Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
by: Paim, Kayua Oleques, et al.
Published: (2025) -
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025) -
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
by: Schaeffer, Joachim, et al.
Published: (2026) -
Toward Autonomous and Efficient Cybersecurity: A Multi-Objective AutoML-based Intrusion Detection System
by: Yang, Li, et al.
Published: (2025)