Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Muzsai, Lajos, Imolai, David, Lukács, András |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024)
by: Muzsai, Lajos, et al.
Published: (2024)
PACT: Reducing Alert Fatigue in Low-Prevalence SOC Streams with Triggered Active Learning
by: Ndichu, Samuel, et al.
Published: (2026)
by: Ndichu, Samuel, et al.
Published: (2026)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
by: Huff, Philip, et al.
Published: (2026)
by: Huff, Philip, et al.
Published: (2026)
Expanding the Attack Scenarios of SAE J1939: A Comprehensive Analysis of Established and Novel Vulnerabilities in Transport Protocol
by: Lee, Hwejae, et al.
Published: (2024)
by: Lee, Hwejae, et al.
Published: (2024)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
by: Piao, Yangheran, et al.
Published: (2025)
by: Piao, Yangheran, et al.
Published: (2025)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Quantum Machine Learning for Cyber-Physical Anomaly Detection in Unmanned Aerial Vehicles: A Leakage-Free Evaluation with Proxy-Audited Feature Sets
by: Paredes, Carlos A. Durán, et al.
Published: (2026)
by: Paredes, Carlos A. Durán, et al.
Published: (2026)
Cross-Domain Malware Detection via Probability-Level Fusion of Lightweight Gradient Boosting Models
by: Mohamed, Omar Khalid Ali
Published: (2025)
by: Mohamed, Omar Khalid Ali
Published: (2025)
Exploratory Analysis of Cyberattack Patterns on E-Commerce Platforms Using Statistical Methods
by: Adeniya, Fatimo Adenike
Published: (2025)
by: Adeniya, Fatimo Adenike
Published: (2025)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
by: Mitchell, Richard Joseph
Published: (2026)
by: Mitchell, Richard Joseph
Published: (2026)
Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
by: Gu, Jintao, et al.
Published: (2025)
by: Gu, Jintao, et al.
Published: (2025)
VOLTRON: Detecting Unknown Malware Using Graph-Based Zero-Shot Learning
by: Akdeniz, M. Tahir, et al.
Published: (2025)
by: Akdeniz, M. Tahir, et al.
Published: (2025)
RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection
by: Buonocore, Tommaso Mario, et al.
Published: (2025)
by: Buonocore, Tommaso Mario, et al.
Published: (2025)
Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives
by: He, Shijing, et al.
Published: (2025)
by: He, Shijing, et al.
Published: (2025)
A Method for Quantifying Human Risk and a Blueprint for LLM Integration
by: Canale, Giuseppe
Published: (2025)
by: Canale, Giuseppe
Published: (2025)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
by: Merves, Tyler H., et al.
Published: (2026)
by: Merves, Tyler H., et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Constant-Size Cryptographic Evidence Structures for Regulated AI Workflows
by: Kao, Leo
Published: (2025)
by: Kao, Leo
Published: (2025)
Control Physiology: An Agent-Based Model of FAIR-CAM Dynamics
by: Jones, Jack, et al.
Published: (2026)
by: Jones, Jack, et al.
Published: (2026)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations
by: Fatouros, George, et al.
Published: (2026)
by: Fatouros, George, et al.
Published: (2026)
Agentic Artificial Intelligence for Ethical Cybersecurity in Uganda: A Reinforcement Learning Framework for Threat Detection in Resource-Constrained Environments
by: Adabara, Ibrahim, et al.
Published: (2025)
by: Adabara, Ibrahim, et al.
Published: (2025)
Safety, Security, and Cognitive Risks in World Models
by: Parmar, Manoj
Published: (2026)
by: Parmar, Manoj
Published: (2026)
Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
by: Engelberg, Gal, et al.
Published: (2026)
by: Engelberg, Gal, et al.
Published: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
DRIFT: Drift-Resilient Invariant-Feature Transformer for DGA Detection
by: Lee, Chaeyoung, et al.
Published: (2026)
by: Lee, Chaeyoung, et al.
Published: (2026)
Developing a Strong CPS Defender: An Evolutionary Approach
by: Hu, Qingyuan, et al.
Published: (2025)
by: Hu, Qingyuan, et al.
Published: (2025)
False Security Confidence in Benign LLM Code Generation
by: Ren, Xiaolei
Published: (2026)
by: Ren, Xiaolei
Published: (2026)
RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code
by: Pellew, John, et al.
Published: (2026)
by: Pellew, John, et al.
Published: (2026)
One-Shot Secure Aggregation: A Hybrid Cryptographic Protocol for Private Federated Learning in IoT
by: Emmaka, Imraul, et al.
Published: (2025)
by: Emmaka, Imraul, et al.
Published: (2025)
Minimizing the Number of Roles in Bottom-Up Role-Mining using Maximal Biclique Enumeration
by: Tripunitara, Mahesh
Published: (2024)
by: Tripunitara, Mahesh
Published: (2024)
Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance
by: Huwyler, Hernan
Published: (2025)
by: Huwyler, Hernan
Published: (2025)
A TEE-Based Architecture for Confidential and Dependable Process Attestation in Authorship Verification
by: Condrey, David
Published: (2026)
by: Condrey, David
Published: (2026)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
by: Filus, Katarzyna, et al.
Published: (2025)
by: Filus, Katarzyna, et al.
Published: (2025)
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
by: Stantchev, Vladimir
Published: (2026)
by: Stantchev, Vladimir
Published: (2026)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
by: Ahi, Kiarash, et al.
Published: (2026)
by: Ahi, Kiarash, et al.
Published: (2026)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
by: Agarwal, Abhinav
Published: (2026)
by: Agarwal, Abhinav
Published: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025)
by: Hoang, Tien Dat
Published: (2025)
SCAFDS: Edge-Feature Graph Attention for Interbank Fraud Detection with Attribution-Grounded SAR Generation
by: Uddin, Mohammad Nasir
Published: (2026)
by: Uddin, Mohammad Nasir
Published: (2026)
Similar Items
-
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024) -
PACT: Reducing Alert Fatigue in Low-Prevalence SOC Streams with Triggered Active Learning
by: Ndichu, Samuel, et al.
Published: (2026) -
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
by: Huff, Philip, et al.
Published: (2026) -
Expanding the Attack Scenarios of SAE J1939: A Comprehensive Analysis of Established and Novel Vulnerabilities in Transport Protocol
by: Lee, Hwejae, et al.
Published: (2024) -
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
by: Piao, Yangheran, et al.
Published: (2025)