Why you shouldn't fully trust ChatGPT: A synthesis of this AI tool's error rates across disciplines and the software engineering lifecycle
Fuente:
arXiv
Saved in:
| Main Author: | Garousi, Vahid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From product to system network challenges in system of systems lifecycle management
by: Salehi, Vahid, et al.
Published: (2025)
by: Salehi, Vahid, et al.
Published: (2025)
Which AI Technique Is Better to Classify Requirements? An Experiment with SVM, LSTM, and ChatGPT
by: El-Hajjami, Abdelkarim, et al.
Published: (2023)
by: El-Hajjami, Abdelkarim, et al.
Published: (2023)
WIP: Assessing the Effectiveness of ChatGPT in Preparatory Testing Activities
by: Haldar, Susmita, et al.
Published: (2025)
by: Haldar, Susmita, et al.
Published: (2025)
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025)
by: Manik, Md Motaleb Hossen
Published: (2025)
Programming with AI: Evaluating ChatGPT, Gemini, AlphaCode, and GitHub Copilot for Programmers
by: Siam, Md Kamrul, et al.
Published: (2024)
by: Siam, Md Kamrul, et al.
Published: (2024)
A pragmatic look at education and training of software test engineers: Further cooperation of academia and industry is needed
by: Garousi, Vahid, et al.
Published: (2024)
by: Garousi, Vahid, et al.
Published: (2024)
Is ChatGPT a Good Software Librarian? An Exploratory Study on the Use of ChatGPT for Software Library Recommendations
by: Latendresse, Jasmine, et al.
Published: (2024)
by: Latendresse, Jasmine, et al.
Published: (2024)
ChatGPT Incorrectness Detection in Software Reviews
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
Exploring Zero-Shot App Review Classification with ChatGPT: Challenges and Potential
by: Chaudhary, Mohit, et al.
Published: (2025)
by: Chaudhary, Mohit, et al.
Published: (2025)
System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT
by: Bhatia, Shreya, et al.
Published: (2024)
by: Bhatia, Shreya, et al.
Published: (2024)
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
by: Kabir, Samia, et al.
Published: (2023)
by: Kabir, Samia, et al.
Published: (2023)
Can ChatGPT support software verification?
by: Janßen, Christian, et al.
Published: (2023)
by: Janßen, Christian, et al.
Published: (2023)
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Evaluation of the Code Generation Capabilities of ChatGPT 4: A Comparative Analysis in 19 Programming Languages
by: Gilbert, L. C.
Published: (2025)
by: Gilbert, L. C.
Published: (2025)
Evaluating ChatGPT-3.5 Efficiency in Solving Coding Problems of Different Complexity Levels: An Empirical Analysis
by: Li, Minda, et al.
Published: (2024)
by: Li, Minda, et al.
Published: (2024)
Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
by: Della Porta, Antonio, et al.
Published: (2025)
by: Della Porta, Antonio, et al.
Published: (2025)
A case study on the transformative potential of AI in software engineering on LeetCode and ChatGPT
by: Merkel, Manuel, et al.
Published: (2025)
by: Merkel, Manuel, et al.
Published: (2025)
Exploring ChatGPT's Capabilities on Vulnerability Management
by: Liu, Peiyu, et al.
Published: (2023)
by: Liu, Peiyu, et al.
Published: (2023)
Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
by: Hyun, Sangwon, et al.
Published: (2025)
by: Hyun, Sangwon, et al.
Published: (2025)
Evaluating Privacy Questions From Stack Overflow: Can ChatGPT Compete?
by: Delile, Zack, et al.
Published: (2023)
by: Delile, Zack, et al.
Published: (2023)
A Qualitative Study on Using ChatGPT for Software Security: Perception vs. Practicality
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
by: Liu, Shigang, et al.
Published: (2024)
by: Liu, Shigang, et al.
Published: (2024)
Are ChatGPT and Other Similar Systems the Modern Lernaean Hydras of AI?
by: Ioannidis, Dimitrios, et al.
Published: (2023)
by: Ioannidis, Dimitrios, et al.
Published: (2023)
Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
by: Arefin, Sayed Erfan, et al.
Published: (2023)
by: Arefin, Sayed Erfan, et al.
Published: (2023)
Benefits and Risks of Using ChatGPT4 as a Teaching Assistant for Computer Science Students
by: Aragonés-Soria, Yaiza, et al.
Published: (2024)
by: Aragonés-Soria, Yaiza, et al.
Published: (2024)
AI-powered software testing tools: A systematic review and empirical assessment of their features and limitations
by: Garousi, Vahid, et al.
Published: (2024)
by: Garousi, Vahid, et al.
Published: (2024)
Why Do Developers Engage with ChatGPT in Issue-Tracker? Investigating Usage and Reliance on ChatGPT-Generated Code
by: Das, Joy Krishan, et al.
Published: (2024)
by: Das, Joy Krishan, et al.
Published: (2024)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Just another copy and paste? Comparing the security vulnerabilities of ChatGPT generated code and StackOverflow answers
by: Hamer, Sivana, et al.
Published: (2024)
by: Hamer, Sivana, et al.
Published: (2024)
The importance of visual modelling languages in generative software engineering
by: Rossi, Roberto
Published: (2024)
by: Rossi, Roberto
Published: (2024)
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation
by: Zhong, Li, et al.
Published: (2023)
by: Zhong, Li, et al.
Published: (2023)
Exploring the Efficacy of Robotic Assistants with ChatGPT and Claude in Enhancing ADHD Therapy: Innovating Treatment Paradigms
by: Berrezueta-Guzman, Santiago, et al.
Published: (2024)
by: Berrezueta-Guzman, Santiago, et al.
Published: (2024)
An Empirical Exploration of ChatGPT's Ability to Support Problem Formulation Tasks for Mission Engineering and a Documentation of its Performance Variability
by: Ofsa, Max, et al.
Published: (2025)
by: Ofsa, Max, et al.
Published: (2025)
Exploring the extent of similarities in software failures across industries using LLMs
by: Detloff, Martin
Published: (2024)
by: Detloff, Martin
Published: (2024)
Safety Analysis in the Era of Large Language Models: A Case Study of STPA using ChatGPT
by: Qi, Yi, et al.
Published: (2023)
by: Qi, Yi, et al.
Published: (2023)
Sentiment analysis for software engineering: How far can zero-shot learning (ZSL) go?
by: Alfayez, Reem, et al.
Published: (2026)
by: Alfayez, Reem, et al.
Published: (2026)
ChatGPT on the Road: Leveraging Large Language Model-Powered In-vehicle Conversational Agents for Safer and More Enjoyable Driving Experience
by: Bond, Yeana Lee, et al.
Published: (2025)
by: Bond, Yeana Lee, et al.
Published: (2025)
The Reproducible Research Platform establishes a unified open science environment bridging data and software lifecycles across disciplines, from proposal to publication
by: Cuny, Andreas P., et al.
Published: (2025)
by: Cuny, Andreas P., et al.
Published: (2025)
Similar Items
-
From product to system network challenges in system of systems lifecycle management
by: Salehi, Vahid, et al.
Published: (2025) -
Which AI Technique Is Better to Classify Requirements? An Experiment with SVM, LSTM, and ChatGPT
by: El-Hajjami, Abdelkarim, et al.
Published: (2023) -
WIP: Assessing the Effectiveness of ChatGPT in Preparatory Testing Activities
by: Haldar, Susmita, et al.
Published: (2025) -
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025) -
Programming with AI: Evaluating ChatGPT, Gemini, AlphaCode, and GitHub Copilot for Programmers
by: Siam, Md Kamrul, et al.
Published: (2024)