Is Your Automated Software Engineer Trustworthy?
Fuente:
arXiv
Saved in:
| Main Authors: | Mathews, Noble Saji, Nagappan, Meiyappan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Design choices made by LLM-based test generators prevent them from finding bugs
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
by: Khatib, Lara, et al.
Published: (2025)
by: Khatib, Lara, et al.
Published: (2025)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
FuzzSlice: Pruning False Positives in Static Analysis Warnings Through Function-Level Fuzzing
by: Murali, Aniruddhan, et al.
Published: (2024)
by: Murali, Aniruddhan, et al.
Published: (2024)
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Whodunit: Classifying Code as Human Authored or GPT-4 Generated -- A case study on CodeChef problems
by: Idialu, Oseremen Joy, et al.
Published: (2024)
by: Idialu, Oseremen Joy, et al.
Published: (2024)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
by: Asare, Owura, et al.
Published: (2022)
by: Asare, Owura, et al.
Published: (2022)
A User-centered Security Evaluation of Copilot
by: Asare, Owura, et al.
Published: (2023)
by: Asare, Owura, et al.
Published: (2023)
RLocator: Reinforcement Learning for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2023)
by: Chakraborty, Partha, et al.
Published: (2023)
Examining LLMs Ability to Summarize Code Through Mutation-Analysis
by: Khatib, Lara, et al.
Published: (2026)
by: Khatib, Lara, et al.
Published: (2026)
Measuring the Runtime Performance of C++ Code Written by Humans using GitHub Copilot
by: Erhabor, Daniel, et al.
Published: (2023)
by: Erhabor, Daniel, et al.
Published: (2023)
BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Towards Understanding Barriers and Mitigation Strategies of Software Engineers with Non-traditional Educational and Occupational Backgrounds
by: Barnes, Tavian, et al.
Published: (2022)
by: Barnes, Tavian, et al.
Published: (2022)
Trustworthy AI Software Engineers
by: Aleti, Aldeida, et al.
Published: (2026)
by: Aleti, Aldeida, et al.
Published: (2026)
Aligning Programming Language and Natural Language: Exploring Design Choices in Multi-Modal Transformer-Based Embedding for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Engineering Trustworthy Software: A Mission for LLMs
by: Vieira, Marco
Published: (2024)
by: Vieira, Marco
Published: (2024)
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View Graphs
by: Mathai, Alex, et al.
Published: (2024)
by: Mathai, Alex, et al.
Published: (2024)
Towards Trustworthy Sentiment Analysis in Software Engineering: Dataset Characteristics and Tool Selection
by: Obaidi, Martin, et al.
Published: (2025)
by: Obaidi, Martin, et al.
Published: (2025)
Automated Quantum Software and AI Engineering
by: Siavash, Nazanin, et al.
Published: (2026)
by: Siavash, Nazanin, et al.
Published: (2026)
AddressWatcher: Sanitizer-Based Localization of Memory Leak Fixes
by: Murali, Aniruddhan, et al.
Published: (2024)
by: Murali, Aniruddhan, et al.
Published: (2024)
TrustOps: Continuously Building Trustworthy Software
by: Brito, Eduardo, et al.
Published: (2024)
by: Brito, Eduardo, et al.
Published: (2024)
Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening
by: Balogh, Gergő, et al.
Published: (2025)
by: Balogh, Gergő, et al.
Published: (2025)
Automated Personnel Selection for Software Engineers Using LLM-Based Profile Evaluation
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
Trustworthy Software Project Generation : a Case Study with an Interactive Theorem Prover
by: Fang, Jian, et al.
Published: (2026)
by: Fang, Jian, et al.
Published: (2026)
Towards Trustworthy AI Software Development Assistance
by: Maninger, Daniel, et al.
Published: (2023)
by: Maninger, Daniel, et al.
Published: (2023)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
A Viable Paradigm of Software Automation: Iterative End-to-End Automated Software Development
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Automated Testing of the GUI of a Real-Life Engineering Software using Large Language Models
by: Rosenbach, Tim, et al.
Published: (2025)
by: Rosenbach, Tim, et al.
Published: (2025)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
by: Mastropaolo, Antonio, et al.
Published: (2025)
by: Mastropaolo, Antonio, et al.
Published: (2025)
Reporting LLM Prompting in Automated Software Engineering: A Guideline Based on Current Practices and Expectations
by: Korn, Alexander, et al.
Published: (2026)
by: Korn, Alexander, et al.
Published: (2026)
Generative Software Engineering
by: Huang, Yuan, et al.
Published: (2024)
by: Huang, Yuan, et al.
Published: (2024)
Do Research Software Engineers and Software Engineering Researchers Speak the Same Language?
by: Kehrer, Timo, et al.
Published: (2025)
by: Kehrer, Timo, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Quantum Software Engineering and Potential of Quantum Computing in Software Engineering Research: A Review
by: Mandal, Ashis Kumar, et al.
Published: (2025)
by: Mandal, Ashis Kumar, et al.
Published: (2025)
Teaching Software Metrology: The Science of Measurement for Software Engineering
by: Ralph, Paul, et al.
Published: (2024)
by: Ralph, Paul, et al.
Published: (2024)
Infusion of Blockchain to Establish Trustworthiness in AI Supported Software Evolution: A Systematic Literature Review
by: Naserameri, Mohammad, et al.
Published: (2026)
by: Naserameri, Mohammad, et al.
Published: (2026)
Automated Trustworthiness Testing for Machine Learning Classifiers
by: Cho, Steven, et al.
Published: (2024)
by: Cho, Steven, et al.
Published: (2024)
Similar Items
-
Design choices made by LLM-based test generators prevent them from finding bugs
by: Mathews, Noble Saji, et al.
Published: (2024) -
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024) -
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
by: Khatib, Lara, et al.
Published: (2025) -
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025) -
FuzzSlice: Pruning False Positives in Static Analysis Warnings Through Function-Level Fuzzing
by: Murali, Aniruddhan, et al.
Published: (2024)