Evaluating AI cyber capabilities with crowdsourced elicitation
Fuente:
arXiv
Saved in:
| Main Authors: | Petrov, Artem, Volkov, Dmitrii |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024)
by: Turtayev, Rustem, et al.
Published: (2024)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025)
by: Reworr, et al.
Published: (2025)
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024)
by: Volkov, Dmitrii
Published: (2024)
AI security and cyber risk in IoT systems
by: Radanliev, Petar, et al.
Published: (2024)
by: Radanliev, Petar, et al.
Published: (2024)
On the use of neurosymbolic AI for defending against cyber attacks
by: Grov, Gudmund, et al.
Published: (2024)
by: Grov, Gudmund, et al.
Published: (2024)
MemTrust: A Zero-Trust Architecture for Unified AI Memory System
by: Zhou, Xing, et al.
Published: (2026)
by: Zhou, Xing, et al.
Published: (2026)
Optimized detection of cyber-attacks on IoT networks via hybrid deep learning models
by: Bensaoud, Ahmed, et al.
Published: (2025)
by: Bensaoud, Ahmed, et al.
Published: (2025)
NEST: Nascent Encoded Steganographic Thoughts
by: Karpov, Artem
Published: (2026)
by: Karpov, Artem
Published: (2026)
Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models
by: Belfiore, Asia, et al.
Published: (2025)
by: Belfiore, Asia, et al.
Published: (2025)
LLM-Guided Prompt Evolution for Password Guessing
by: Mazin, Vladimir A., et al.
Published: (2026)
by: Mazin, Vladimir A., et al.
Published: (2026)
CyberRAG: An Agentic RAG cyber attack classification and reporting tool
by: Blefari, Francesco, et al.
Published: (2025)
by: Blefari, Francesco, et al.
Published: (2025)
Technical Evaluation of a Disruptive Approach in Homomorphic AI
by: Filiol, Eric
Published: (2025)
by: Filiol, Eric
Published: (2025)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025)
by: Rodriguez, Mikel, et al.
Published: (2025)
Referential Security as a New Paradigm for AI Evaluations
by: Ristea, Dan, et al.
Published: (2026)
by: Ristea, Dan, et al.
Published: (2026)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents
by: Aonzo, Simone, et al.
Published: (2026)
by: Aonzo, Simone, et al.
Published: (2026)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art
by: Trinh, Khoi, et al.
Published: (2026)
by: Trinh, Khoi, et al.
Published: (2026)
VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces
by: Grigor, Artem, et al.
Published: (2025)
by: Grigor, Artem, et al.
Published: (2025)
SPEAR: Security Posture Evaluation using AI Planner-Reasoning on Attack-Connectivity Hypergraphs
by: Podder, Rakesh, et al.
Published: (2025)
by: Podder, Rakesh, et al.
Published: (2025)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
by: Li, Changyi, et al.
Published: (2026)
by: Li, Changyi, et al.
Published: (2026)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
by: Usynin, Dmitrii, et al.
Published: (2023)
by: Usynin, Dmitrii, et al.
Published: (2023)
A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
by: Wang, Jinghao, et al.
Published: (2025)
by: Wang, Jinghao, et al.
Published: (2025)
Malware analysis assisted by AI with R2AI
by: Apvrille, Axelle, et al.
Published: (2025)
by: Apvrille, Axelle, et al.
Published: (2025)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
AI Identity: Standards, Gaps, and Research Directions for AI Agents
by: Otsuka, Takumi, et al.
Published: (2026)
by: Otsuka, Takumi, et al.
Published: (2026)
NetMoniAI: An Agentic AI Framework for Network Security & Monitoring
by: Zambare, Pallavi, et al.
Published: (2025)
by: Zambare, Pallavi, et al.
Published: (2025)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI
by: Huang, Ken, et al.
Published: (2025)
by: Huang, Ken, et al.
Published: (2025)
AI Security Map: Holistic Organization of AI Security Technologies and Impacts on Stakeholders
by: Kato, Hiroya, et al.
Published: (2025)
by: Kato, Hiroya, et al.
Published: (2025)
STRIDE-AI: A Threat Modeling Framework for Generative AI Security Assessment
by: Cyrille, Tsafac Nkombong Regine, et al.
Published: (2026)
by: Cyrille, Tsafac Nkombong Regine, et al.
Published: (2026)
BadGPT-4o: stripping safety finetuning from GPT models
by: Krupkina, Ekaterina, et al.
Published: (2024)
by: Krupkina, Ekaterina, et al.
Published: (2024)
Integrative Approaches in Cybersecurity and AI
by: Omar, Marwan
Published: (2024)
by: Omar, Marwan
Published: (2024)
Security of and by Generative AI platforms
by: Hayagreevan, Hari, et al.
Published: (2024)
by: Hayagreevan, Hari, et al.
Published: (2024)
AI Native Asset Intelligence
by: Engelberg, Gal, et al.
Published: (2026)
by: Engelberg, Gal, et al.
Published: (2026)
Similar Items
-
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024) -
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024) -
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025) -
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024) -
AI security and cyber risk in IoT systems
by: Radanliev, Petar, et al.
Published: (2024)