LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Camporese, Maria, Massacci, Fabio, Gong, Yuanjun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs
by: Camporese, Maria, et al.
Published: (2025)
by: Camporese, Maria, et al.
Published: (2025)
Using ML filters to help automated vulnerability repairs: when it helps and when it doesn't
by: Camporese, Maria, et al.
Published: (2025)
by: Camporese, Maria, et al.
Published: (2025)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025)
by: Roig, JV
Published: (2025)
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
by: Shah, Syed Tauhid Ullah, et al.
Published: (2025)
by: Shah, Syed Tauhid Ullah, et al.
Published: (2025)
Analyzing and Mitigating (with LLMs) the Security Misconfigurations of Helm Charts from Artifact Hub
by: Minna, Francesco, et al.
Published: (2024)
by: Minna, Francesco, et al.
Published: (2024)
Risks of ignoring uncertainty propagation in AI-augmented security pipelines
by: Mezzi, Emanuele, et al.
Published: (2024)
by: Mezzi, Emanuele, et al.
Published: (2024)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Analysis of LLMs vs Human Experts in Requirements Engineering
by: Hymel, Cory, et al.
Published: (2025)
by: Hymel, Cory, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
A3Rank: Augmentation Alignment Analysis for Prioritizing Overconfident Failing Samples for Deep Learning Models
by: Wei, Zhengyuan, et al.
Published: (2024)
by: Wei, Zhengyuan, et al.
Published: (2024)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
On the Abolition of the "ICSE Paper" and the Adoption of the "Registered Proposal" and the "Results Report"
by: Massacci, Fabio, et al.
Published: (2026)
by: Massacci, Fabio, et al.
Published: (2026)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
Application Modernization with LLMs: Addressing Core Challenges in Reliability, Security, and Quality
by: Ponnusamy, Ahilan Ayyachamy Nadar
Published: (2025)
by: Ponnusamy, Ahilan Ayyachamy Nadar
Published: (2025)
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors
by: Yu, Jiaxin, et al.
Published: (2024)
by: Yu, Jiaxin, et al.
Published: (2024)
LLMs + Security = Trouble
by: Livshits, Benjamin
Published: (2026)
by: Livshits, Benjamin
Published: (2026)
A Systematic Literature Review on Automated Exploit and Security Test Generation
by: Bui, Quang-Cuong, et al.
Published: (2025)
by: Bui, Quang-Cuong, et al.
Published: (2025)
Can LLMs Replace Humans During Code Chunking?
by: Glasz, Christopher, et al.
Published: (2025)
by: Glasz, Christopher, et al.
Published: (2025)
Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
Five Fatal Assumptions: Why T-Shirt Sizing Systematically Fails for AI Projects
by: Soundaramourty, Raja, et al.
Published: (2026)
by: Soundaramourty, Raja, et al.
Published: (2026)
Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
by: Jahan, Sigma, et al.
Published: (2025)
by: Jahan, Sigma, et al.
Published: (2025)
Vendor-Aware Industrial Agents: RAG-Enhanced LLMs for Secure On-Premise PLC Code Generation
by: Kersting, Joschka, et al.
Published: (2025)
by: Kersting, Joschka, et al.
Published: (2025)
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
by: Kim, Myeongsoo, et al.
Published: (2026)
by: Kim, Myeongsoo, et al.
Published: (2026)
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
by: Lu, Ruofan, et al.
Published: (2025)
by: Lu, Ruofan, et al.
Published: (2025)
Analysis on LLMs Performance for Code Summarization
by: Akib, Md. Ahnaf, et al.
Published: (2024)
by: Akib, Md. Ahnaf, et al.
Published: (2024)
MESIA: Understanding and Leveraging Supplementary Nature of Method-level Comments for Automatic Comment Generation
by: Pan, Xinglu, et al.
Published: (2024)
by: Pan, Xinglu, et al.
Published: (2024)
Mastering the Craft of Data Synthesis for CodeLLMs
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
by: Pan, Chenkai, et al.
Published: (2026)
by: Pan, Chenkai, et al.
Published: (2026)
Impact of Comments on LLM Comprehension of Legacy Code
by: Sabetto, Rock, et al.
Published: (2025)
by: Sabetto, Rock, et al.
Published: (2025)
LLMs for Science: Usage for Code Generation and Data Analysis
by: Nejjar, Mohamed, et al.
Published: (2023)
by: Nejjar, Mohamed, et al.
Published: (2023)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Leveraging Metamemory Mechanisms for Enhanced Data-Free Code Generation in LLMs
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
A Qualitative Study on Using ChatGPT for Software Security: Perception vs. Practicality
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
by: Trinkenreich, Bianca, et al.
Published: (2025)
by: Trinkenreich, Bianca, et al.
Published: (2025)
The Energy Footprint of LLM-Based Environmental Analysis: LLMs and Domain Products
by: Bao, Alicia, et al.
Published: (2026)
by: Bao, Alicia, et al.
Published: (2026)
The Hitchhiker's Guide to Program Analysis, Part II: Deep Thoughts by LLMs
by: Li, Haonan, et al.
Published: (2025)
by: Li, Haonan, et al.
Published: (2025)
Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
by: Yamani, Asma, et al.
Published: (2025)
by: Yamani, Asma, et al.
Published: (2025)
TAAF: A Trace Abstraction and Analysis Framework Synergizing Knowledge Graphs and LLMs
by: Ezaz, Alireza, et al.
Published: (2026)
by: Ezaz, Alireza, et al.
Published: (2026)
Revisiting the Role of Natural Language Code Comments in Code Translation
by: Gupta, Monika, et al.
Published: (2026)
by: Gupta, Monika, et al.
Published: (2026)
Similar Items
-
Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs
by: Camporese, Maria, et al.
Published: (2025) -
Using ML filters to help automated vulnerability repairs: when it helps and when it doesn't
by: Camporese, Maria, et al.
Published: (2025) -
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025) -
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
by: Shah, Syed Tauhid Ullah, et al.
Published: (2025) -
Analyzing and Mitigating (with LLMs) the Security Misconfigurations of Helm Charts from Artifact Hub
by: Minna, Francesco, et al.
Published: (2024)