Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
Fuente:
arXiv
Saved in:
| Main Authors: | Sadik, Ahmed R., Govind, Siddhata |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025)
by: Manik, Md Motaleb Hossen
Published: (2025)
Zero-shot Evaluation of Deep Learning for Java Code Clone Detection
by: Heinze, Thomas S.
Published: (2026)
by: Heinze, Thomas S.
Published: (2026)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
by: Guo, Daya, et al.
Published: (2024)
by: Guo, Daya, et al.
Published: (2024)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
by: Khan, M Zafir Sadik, et al.
Published: (2026)
by: Khan, M Zafir Sadik, et al.
Published: (2026)
Assessing GPT-4-Vision's Capabilities in UML-Based Code Generation
by: Antal, Gábor, et al.
Published: (2024)
by: Antal, Gábor, et al.
Published: (2024)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Defusing Logic Bombs in Symbolic Execution with LLM-Generated Ghost Code
by: Bouras, Dimitrios Stamatios, et al.
Published: (2026)
by: Bouras, Dimitrios Stamatios, et al.
Published: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
RacerF: Lightweight Static Data Race Detection for C Code
by: Dacík, Tomáš, et al.
Published: (2025)
by: Dacík, Tomáš, et al.
Published: (2025)
How Natural Language Proficiency Shapes GenAI Code for Software Engineering Tasks
by: Rojpaisarnkit, Ruksit, et al.
Published: (2025)
by: Rojpaisarnkit, Ruksit, et al.
Published: (2025)
On Repairing Quantum Programs Using ChatGPT
by: Guo, Xiaoyu, et al.
Published: (2024)
by: Guo, Xiaoyu, et al.
Published: (2024)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
by: Petrukha, Ivan, et al.
Published: (2025)
by: Petrukha, Ivan, et al.
Published: (2025)
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
by: Misra, Diganta, et al.
Published: (2025)
by: Misra, Diganta, et al.
Published: (2025)
Neural Code Translation of Legacy Code: APL to C#
by: Ramadan, Abdulrahman, et al.
Published: (2026)
by: Ramadan, Abdulrahman, et al.
Published: (2026)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
Building A Proof-Oriented Programmer That Is 64% Better Than GPT-4o Under Data Scarcity
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Dynamic Stability of LLM-Generated Code
by: Rajput, Prateek, et al.
Published: (2025)
by: Rajput, Prateek, et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
Specification and Detection of LLM Code Smells
by: Mahmoudi, Brahim, et al.
Published: (2025)
by: Mahmoudi, Brahim, et al.
Published: (2025)
Effective LLM-Driven Code Generation with Pythoness
by: Levin, Kyla H., et al.
Published: (2025)
by: Levin, Kyla H., et al.
Published: (2025)
AI-Mediated Code Comment Improvement
by: Dhakal, Maria, et al.
Published: (2025)
by: Dhakal, Maria, et al.
Published: (2025)
Micro-Patterns in Solidity Code
by: Ruschioni, Luca, et al.
Published: (2025)
by: Ruschioni, Luca, et al.
Published: (2025)
Pareto Optimal Code Generation
by: Orlanski, Gabriel, et al.
Published: (2025)
by: Orlanski, Gabriel, et al.
Published: (2025)
SALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary Decompilation
by: Wang, Yongpan, et al.
Published: (2025)
by: Wang, Yongpan, et al.
Published: (2025)
CodeFuse-Query: A Data-Centric Static Code Analysis System for Large-Scale Organizations
by: Xie, Xiaoheng, et al.
Published: (2024)
by: Xie, Xiaoheng, et al.
Published: (2024)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
by: Dandamudi, Rohit, et al.
Published: (2024)
by: Dandamudi, Rohit, et al.
Published: (2024)
o3-mini vs DeepSeek-R1: Which One is Safer?
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
Automated Code Editing with Search-Generate-Modify
by: Liu, Changshu, et al.
Published: (2023)
by: Liu, Changshu, et al.
Published: (2023)
Validated Code Translation for Projects with External Libraries
by: Zhang, Hanliang, et al.
Published: (2026)
by: Zhang, Hanliang, et al.
Published: (2026)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
A Trace-based Approach for Code Safety Analysis
by: Xu, Hui
Published: (2025)
by: Xu, Hui
Published: (2025)
Similar Items
-
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025) -
Zero-shot Evaluation of Deep Learning for Java Code Clone Detection
by: Heinze, Thomas S.
Published: (2026) -
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
by: Guo, Daya, et al.
Published: (2024) -
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025) -
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)