Are Large Language Models a Threat to Programming Platforms? An Exploratory Study

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Billah, Md Mustakim, Roy, Palash Ranjan, Codabux, Zadia, Roy, Banani
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917771810439168
author Billah, Md Mustakim
Roy, Palash Ranjan
Codabux, Zadia
Roy, Banani
author_facet Billah, Md Mustakim
Roy, Palash Ranjan
Codabux, Zadia
Roy, Banani
contents Competitive programming platforms like LeetCode, Codeforces, and HackerRank evaluate programming skills, often used by recruiters for screening. With the rise of advanced Large Language Models (LLMs) such as ChatGPT, Gemini, and Meta AI, their problem-solving ability on these platforms needs assessment. This study explores LLMs' ability to tackle diverse programming challenges across platforms with varying difficulty, offering insights into their real-time and offline performance and comparing them with human programmers. We tested 98 problems from LeetCode, 126 from Codeforces, covering 15 categories. Nine online contests from Codeforces and LeetCode were conducted, along with two certification tests on HackerRank, to assess real-time performance. Prompts and feedback mechanisms were used to guide LLMs, and correlations were explored across different scenarios. LLMs, like ChatGPT (71.43% success on LeetCode), excelled in LeetCode and HackerRank certifications but struggled in virtual contests, particularly on Codeforces. They performed better than users in LeetCode archives, excelling in time and memory efficiency but underperforming in harder Codeforces contests. While not immediately threatening, LLMs performance on these platforms is concerning, and future improvements will need addressing.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Large Language Models a Threat to Programming Platforms? An Exploratory Study
Billah, Md Mustakim
Roy, Palash Ranjan
Codabux, Zadia
Roy, Banani
Software Engineering
Competitive programming platforms like LeetCode, Codeforces, and HackerRank evaluate programming skills, often used by recruiters for screening. With the rise of advanced Large Language Models (LLMs) such as ChatGPT, Gemini, and Meta AI, their problem-solving ability on these platforms needs assessment. This study explores LLMs' ability to tackle diverse programming challenges across platforms with varying difficulty, offering insights into their real-time and offline performance and comparing them with human programmers. We tested 98 problems from LeetCode, 126 from Codeforces, covering 15 categories. Nine online contests from Codeforces and LeetCode were conducted, along with two certification tests on HackerRank, to assess real-time performance. Prompts and feedback mechanisms were used to guide LLMs, and correlations were explored across different scenarios. LLMs, like ChatGPT (71.43% success on LeetCode), excelled in LeetCode and HackerRank certifications but struggled in virtual contests, particularly on Codeforces. They performed better than users in LeetCode archives, excelling in time and memory efficiency but underperforming in harder Codeforces contests. While not immediately threatening, LLMs performance on these platforms is concerning, and future improvements will need addressing.
title Are Large Language Models a Threat to Programming Platforms? An Exploratory Study
topic Software Engineering
url https://arxiv.org/abs/2409.05824