Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
Fuente:
arXiv
Guardado en:
| Autores principales: | Honarvar, Shahin, Gorzynski, Amber, Lee-Jones, James, Coppock, Harry, Rei, Marek, Ryan, Joseph, Donaldson, Alastair F. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code
por: Honarvar, Shahin, et al.
Publicado: (2023)
por: Honarvar, Shahin, et al.
Publicado: (2023)
GitOps for Capture the Flag Platforms
por: Albrechtsen, Mikkel Bengtson, et al.
Publicado: (2026)
por: Albrechtsen, Mikkel Bengtson, et al.
Publicado: (2026)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
por: Agrawal, Aryan, et al.
Publicado: (2025)
por: Agrawal, Aryan, et al.
Publicado: (2025)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
por: Al-Kaswan, Ali, et al.
Publicado: (2026)
por: Al-Kaswan, Ali, et al.
Publicado: (2026)
Auto-SPT: Automating Semantic Preserving Transformations for Code
por: Hooda, Ashish, et al.
Publicado: (2025)
por: Hooda, Ashish, et al.
Publicado: (2025)
Capturing the Effects of Quantization on Trojans in Code LLMs
por: Hussain, Aftab, et al.
Publicado: (2025)
por: Hussain, Aftab, et al.
Publicado: (2025)
An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
por: Jia, Nan, et al.
Publicado: (2025)
por: Jia, Nan, et al.
Publicado: (2025)
ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
por: Fu, Lei, et al.
Publicado: (2025)
por: Fu, Lei, et al.
Publicado: (2025)
Sustainability Flags for the Identification of Sustainability Posts in Q&A Platforms
por: Ahmadisakha, Sahar, et al.
Publicado: (2025)
por: Ahmadisakha, Sahar, et al.
Publicado: (2025)
Capturing Semantic Flow of ML-based Systems
por: Yoo, Shin, et al.
Publicado: (2025)
por: Yoo, Shin, et al.
Publicado: (2025)
Artisan: Agentic Artifact Evaluation
por: Baek, Doehyun, et al.
Publicado: (2026)
por: Baek, Doehyun, et al.
Publicado: (2026)
Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution
por: Feng, Rong, et al.
Publicado: (2025)
por: Feng, Rong, et al.
Publicado: (2025)
Simplicity by Obfuscation: Evaluating LLM-Driven Code Transformation with Semantic Elasticity
por: De Tomasi, Lorenzo, et al.
Publicado: (2025)
por: De Tomasi, Lorenzo, et al.
Publicado: (2025)
The Devil Is in the Command Line: Associating the Compiler Flags With the Binary and Build Metadata
por: Kudrjavets, Gunnar, et al.
Publicado: (2023)
por: Kudrjavets, Gunnar, et al.
Publicado: (2023)
A Shallow Embedding of Datalog in Lean
por: Shahin, Ramy
Publicado: (2026)
por: Shahin, Ramy
Publicado: (2026)
Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs
por: CodeArts Model Team, et al.
Publicado: (2026)
por: CodeArts Model Team, et al.
Publicado: (2026)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
por: Roig, JV
Publicado: (2025)
por: Roig, JV
Publicado: (2025)
An Agentic Approach Towards Replication Package Quality Evaluation
por: Mbida, Maximilian Alexander Amougou, et al.
Publicado: (2026)
por: Mbida, Maximilian Alexander Amougou, et al.
Publicado: (2026)
Agentic Repository Mining: A Multi-Task Evaluation
por: Härtel, Johannes
Publicado: (2026)
por: Härtel, Johannes
Publicado: (2026)
Algorithm-Based Pipeline for Reliable and Intent-Preserving Code Translation with LLMs
por: Dipto, Shahriar Rumi, et al.
Publicado: (2026)
por: Dipto, Shahriar Rumi, et al.
Publicado: (2026)
Consistent or Sensitive? Automated Code Revision Tools Against Semantics-Preserving Perturbations
por: Pirouzkhah, Shirin, et al.
Publicado: (2026)
por: Pirouzkhah, Shirin, et al.
Publicado: (2026)
SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System Testing
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
por: Hu, Haichuan, et al.
Publicado: (2024)
por: Hu, Haichuan, et al.
Publicado: (2024)
Semantic-Preserving Transformations as Mutation Operators: A Study on Their Effectiveness in Defect Detection
por: Hort, Max, et al.
Publicado: (2025)
por: Hort, Max, et al.
Publicado: (2025)
Extremal Testing for Network Software using LLMs
por: Singha, Rathin, et al.
Publicado: (2025)
por: Singha, Rathin, et al.
Publicado: (2025)
Behaviour Driven Development Scenario Generation with Large Language Models
por: Rathnayake, Amila, et al.
Publicado: (2026)
por: Rathnayake, Amila, et al.
Publicado: (2026)
One-Year Internship Program on Software Engineering: Students' Perceptions and Educators' Lessons Learned
por: Abaei, Golnoush, et al.
Publicado: (2026)
por: Abaei, Golnoush, et al.
Publicado: (2026)
Using LLMs in Generating Design Rationale for Software Architecture Decisions
por: Zhou, Xiyu, et al.
Publicado: (2025)
por: Zhou, Xiyu, et al.
Publicado: (2025)
Agentic LLMs for REST API Test Amplification: A Comparative Study Across Cloud Applications
por: Besjes, Jarne, et al.
Publicado: (2025)
por: Besjes, Jarne, et al.
Publicado: (2025)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
por: Hasan, Alif Al, et al.
Publicado: (2026)
por: Hasan, Alif Al, et al.
Publicado: (2026)
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors
por: Yu, Jiaxin, et al.
Publicado: (2024)
por: Yu, Jiaxin, et al.
Publicado: (2024)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
por: Sehgal, Atharva, et al.
Publicado: (2026)
por: Sehgal, Atharva, et al.
Publicado: (2026)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
por: Le-Cong, Thanh, et al.
Publicado: (2025)
por: Le-Cong, Thanh, et al.
Publicado: (2025)
Reusing Legacy Code in WebAssembly: Key Challenges of Cross-Compilation and Code Semantics Preservation
por: Baradaran, Sara, et al.
Publicado: (2024)
por: Baradaran, Sara, et al.
Publicado: (2024)
Evaluating Tool Cloning in Agentic-AI Ecosystems
por: Kim, Taein, et al.
Publicado: (2026)
por: Kim, Taein, et al.
Publicado: (2026)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
por: Yang, Hua, et al.
Publicado: (2025)
por: Yang, Hua, et al.
Publicado: (2025)
AgenticAKM : Enroute to Agentic Architecture Knowledge Management
por: Dhar, Rudra, et al.
Publicado: (2026)
por: Dhar, Rudra, et al.
Publicado: (2026)
SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation
por: Radanliev, Petar, et al.
Publicado: (2026)
por: Radanliev, Petar, et al.
Publicado: (2026)
LLMs for Test Input Generation for Semantic Caches
por: Rasool, Zafaryab, et al.
Publicado: (2024)
por: Rasool, Zafaryab, et al.
Publicado: (2024)
AOCI: Symbolic-Semantic Indexing for Practical Repository-Scale Code Understanding with LLMs
por: Liu, Jinshi, et al.
Publicado: (2026)
por: Liu, Jinshi, et al.
Publicado: (2026)
Ejemplares similares
-
Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code
por: Honarvar, Shahin, et al.
Publicado: (2023) -
GitOps for Capture the Flag Platforms
por: Albrechtsen, Mikkel Bengtson, et al.
Publicado: (2026) -
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
por: Agrawal, Aryan, et al.
Publicado: (2025) -
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
por: Al-Kaswan, Ali, et al.
Publicado: (2026) -
Auto-SPT: Automating Semantic Preserving Transformations for Code
por: Hooda, Ashish, et al.
Publicado: (2025)