An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?
Fuente:
arXiv
Saved in:
| Main Authors: | Suh, Hyunjae, Tafreshipour, Mahan, Li, Jiawei, Bhattiprolu, Adithya, Ahmed, Iftekhar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human or LLM? A Comparative Study on Accessible Code Generation Capability
by: Suh, Hyunjae, et al.
Published: (2025)
by: Suh, Hyunjae, et al.
Published: (2025)
Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories
by: Tafreshipour, Mahan, et al.
Published: (2024)
by: Tafreshipour, Mahan, et al.
Published: (2024)
Test Smell: A Parasitic Energy Consumer in Software Testing
by: Misu, Md Rakib Hossain, et al.
Published: (2023)
by: Misu, Md Rakib Hossain, et al.
Published: (2023)
Is Multi-Agent Debate (MAD) the Silver Bullet? An Empirical Analysis of MAD in Code Summarization and Translation
by: Chun, Jina, et al.
Published: (2025)
by: Chun, Jina, et al.
Published: (2025)
Does the Order of Fine-tuning Matter and Why?
by: Chen, Qihong, et al.
Published: (2024)
by: Chen, Qihong, et al.
Published: (2024)
Does Documentation Matter? An Empirical Study of Practitioners' Perspective on Open-Source Software Adoption
by: Imani, Aaron, et al.
Published: (2024)
by: Imani, Aaron, et al.
Published: (2024)
Automatically Detecting Checked-In Secrets in Android Apps: How Far Are We?
by: Li, Kevin, et al.
Published: (2024)
by: Li, Kevin, et al.
Published: (2024)
Can AI Agents Generate Microservices? How Far are We?
by: Adnan, Bassam, et al.
Published: (2026)
by: Adnan, Bassam, et al.
Published: (2026)
Vulnerability Detection with Code Language Models: How Far Are We?
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality
by: Ghammam, Anwar, et al.
Published: (2026)
by: Ghammam, Anwar, et al.
Published: (2026)
Duplicate Bug Report Detection: How Far Are We?
by: Zhang, Ting, et al.
Published: (2022)
by: Zhang, Ting, et al.
Published: (2022)
Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message
by: Li, Jiawei, et al.
Published: (2025)
by: Li, Jiawei, et al.
Published: (2025)
Model Editing for LLMs4Code: How Far are We?
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++
by: Nguyen, Anh The, et al.
Published: (2024)
by: Nguyen, Anh The, et al.
Published: (2024)
Unraveling the Potential of Large Language Models in Code Translation: How Far Are We?
by: Tao, Qingxiao, et al.
Published: (2024)
by: Tao, Qingxiao, et al.
Published: (2024)
Large Language Models for Equivalent Mutant Detection: How Far Are We?
by: Tian, Zhao, et al.
Published: (2024)
by: Tian, Zhao, et al.
Published: (2024)
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
by: Saha, Soumit Kanti, et al.
Published: (2024)
by: Saha, Soumit Kanti, et al.
Published: (2024)
Consider What Humans Consider: Optimizing Commit Message Leveraging Contexts Considered By Human
by: Li, Jiawei, et al.
Published: (2025)
by: Li, Jiawei, et al.
Published: (2025)
Vulnerability-Affected Versions Identification: How Far Are We?
by: Chen, Xingchu, et al.
Published: (2025)
by: Chen, Xingchu, et al.
Published: (2025)
From Bias To Improved Prompts: A Case Study of Bias Mitigation of Clone Detection Models
by: Chen, QiHong, et al.
Published: (2025)
by: Chen, QiHong, et al.
Published: (2025)
Automatically Recommend Code Updates: Are We There Yet?
by: Liu, Yue, et al.
Published: (2022)
by: Liu, Yue, et al.
Published: (2022)
Do AI Coding Agents Log Like Humans? An Empirical Study
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or Worse
by: Imani, Aaron, et al.
Published: (2025)
by: Imani, Aaron, et al.
Published: (2025)
Retrieval-Augmented Test Generation: How Far Are We?
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
When ChatGPT Meets Smart Contract Vulnerability Detection: How Far Are We?
by: Chen, Chong, et al.
Published: (2023)
by: Chen, Chong, et al.
Published: (2023)
How Far Are We? The Triumphs and Trials of Generative AI in Learning Software Engineering
by: Choudhuri, Rudrajit, et al.
Published: (2023)
by: Choudhuri, Rudrajit, et al.
Published: (2023)
Leveraging Language Models for Log Statement Generation in Multilingual Scenarios: How Far Are We?
by: Kusama, Kazuki, et al.
Published: (2026)
by: Kusama, Kazuki, et al.
Published: (2026)
Investigating the Impact of Code Comment Inconsistency on Bug Introducing
by: Radmanesh, Shiva, et al.
Published: (2024)
by: Radmanesh, Shiva, et al.
Published: (2024)
Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study
by: Santana Jr, E. G., et al.
Published: (2025)
by: Santana Jr, E. G., et al.
Published: (2025)
Representation Learning for Stack Overflow Posts: How Far are We?
by: He, Junda, et al.
Published: (2023)
by: He, Junda, et al.
Published: (2023)
Automated Testing of Task-based Chatbots: How Far Are We?
by: Clerissi, Diego, et al.
Published: (2026)
by: Clerissi, Diego, et al.
Published: (2026)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
How Far Have We Gone in Binary Code Understanding Using Large Language Models
by: Shang, Xiuwei, et al.
Published: (2024)
by: Shang, Xiuwei, et al.
Published: (2024)
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?
by: Chen, QiHong, et al.
Published: (2024)
by: Chen, QiHong, et al.
Published: (2024)
Context Conquers Parameters: Outperforming Proprietary LLM in Commit Message Generation
by: Imani, Aaron, et al.
Published: (2024)
by: Imani, Aaron, et al.
Published: (2024)
Automatic Data Labeling for Software Vulnerability Prediction Models: How Far Are We?
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
by: Jiang, Zhihan, et al.
Published: (2023)
by: Jiang, Zhihan, et al.
Published: (2023)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
by: O'Brien, Conor, et al.
Published: (2024)
by: O'Brien, Conor, et al.
Published: (2024)
Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChat
by: Ji, Kexing, et al.
Published: (2025)
by: Ji, Kexing, et al.
Published: (2025)
Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We?
by: Li, Kaixuan, et al.
Published: (2024)
by: Li, Kaixuan, et al.
Published: (2024)
Similar Items
-
Human or LLM? A Comparative Study on Accessible Code Generation Capability
by: Suh, Hyunjae, et al.
Published: (2025) -
Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories
by: Tafreshipour, Mahan, et al.
Published: (2024) -
Test Smell: A Parasitic Energy Consumer in Software Testing
by: Misu, Md Rakib Hossain, et al.
Published: (2023) -
Is Multi-Agent Debate (MAD) the Silver Bullet? An Empirical Analysis of MAD in Code Summarization and Translation
by: Chun, Jina, et al.
Published: (2025) -
Does the Order of Fine-tuning Matter and Why?
by: Chen, Qihong, et al.
Published: (2024)