Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Al-Kaswan, Ali, Plotnikov, Maksim, Hájek, Maxim, Vízner, Roland, van Deursen, Arie, Izadi, Maliheh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025)
by: Al-Kaswan, Ali, et al.
Published: (2025)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024)
by: de Moor, Aral, et al.
Published: (2024)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024)
by: Katzy, Jonathan, et al.
Published: (2024)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks before and after Fine-tuning
by: Salerno, Fabio, et al.
Published: (2025)
by: Salerno, Fabio, et al.
Published: (2025)
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024)
by: Izadi, Maliheh, et al.
Published: (2024)
Evaluating Non-English Developer Support in Machine Learning for Software Engineering
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
HyperSeq: A Hyper-Adaptive Representation for Predictive Sequencing of States
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
by: Li, Ziyou, et al.
Published: (2025)
by: Li, Ziyou, et al.
Published: (2025)
Towards Automatic Translation of Machine Learning Visual Insights to Analytical Assertions
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
A Multi-agent Onboarding Assistant based on Large Language Models, Retrieval Augmented Generation, and Chain-of-Thought
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
by: Mundhra, Yash, et al.
Published: (2025)
by: Mundhra, Yash, et al.
Published: (2025)
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
GitOps for Capture the Flag Platforms
by: Albrechtsen, Mikkel Bengtson, et al.
Published: (2026)
by: Albrechtsen, Mikkel Bengtson, et al.
Published: (2026)
Understanding Feedback Mechanisms in Machine Learning Jupyter Notebooks
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
Data vs. Model Machine Learning Fairness Testing: An Empirical Study
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
by: Deljouyi, Amirhossein, et al.
Published: (2024)
by: Deljouyi, Amirhossein, et al.
Published: (2024)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
by: Popescu, Razvan Mihai, et al.
Published: (2026)
by: Popescu, Razvan Mihai, et al.
Published: (2026)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
by: Kuo, Nadine, et al.
Published: (2026)
by: Kuo, Nadine, et al.
Published: (2026)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
McUDI: Model-Centric Unsupervised Degradation Indicator for Failure Prediction AIOps Solutions
by: Poenaru-Olaru, Lorena, et al.
Published: (2024)
by: Poenaru-Olaru, Lorena, et al.
Published: (2024)
Human-AI Experience in Integrated Development Environments: A Systematic Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2025)
by: Sergeyuk, Agnia, et al.
Published: (2025)
Developer Needs and Feasible Features for AI Assistants in IDEs
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Can Coding Agents Be General Agents?
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
Is Your Anomaly Detector Ready for Change? Adapting AIOps Solutions to the Real World
by: Poenaru-Olaru, Lorena, et al.
Published: (2023)
by: Poenaru-Olaru, Lorena, et al.
Published: (2023)
Exploring LLM-based Agents for Root Cause Analysis
by: Roy, Devjeet, et al.
Published: (2024)
by: Roy, Devjeet, et al.
Published: (2024)
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Code4MeV2: a Research-oriented Code-completion Platform
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
What Challenges Do Developers Face in AI Agent Systems? An Empirical Study on Stack Overflow & GitHub Issues
by: Asgari, Ali, et al.
Published: (2025)
by: Asgari, Ali, et al.
Published: (2025)
Sustainable Machine Learning Retraining: Optimizing Energy Efficiency Without Compromising Accuracy
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
Multi-Agent Systems for Root Cause Analysis in Microservices
by: Naakka, Alexander, et al.
Published: (2026)
by: Naakka, Alexander, et al.
Published: (2026)
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
by: Honarvar, Shahin, et al.
Published: (2026)
by: Honarvar, Shahin, et al.
Published: (2026)
Similar Items
-
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023) -
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026) -
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025) -
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024) -
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)