What You See Is Not Always What You Get: Evaluating GPT's Comprehension of Source Code
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Jiawen, Zhu, Bangshuo, Chen, Huaming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Your Code Agent Can Grow Alongside You with Structured Memory
by: Deng, Yi-Xuan, et al.
Published: (2026)
by: Deng, Yi-Xuan, et al.
Published: (2026)
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
by: Spracklen, Joseph, et al.
Published: (2024)
by: Spracklen, Joseph, et al.
Published: (2024)
Let the Code LLM Edit Itself When You Edit the Code
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
by: Huang, Linghan, et al.
Published: (2025)
by: Huang, Linghan, et al.
Published: (2025)
What You See Is What You Get: Attention-based Self-guided Automatic Unit Test Generation
by: Yin, Xin, et al.
Published: (2024)
by: Yin, Xin, et al.
Published: (2024)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
by: Wu, Yixi, et al.
Published: (2024)
by: Wu, Yixi, et al.
Published: (2024)
Supersonic: Learning to Generate Source Code Optimizations in C/C++
by: Chen, Zimin, et al.
Published: (2023)
by: Chen, Zimin, et al.
Published: (2023)
What You Need is What You Get: Theory of Mind for an LLM-Based Code Understanding Assistant
by: Richards, Jonan, et al.
Published: (2024)
by: Richards, Jonan, et al.
Published: (2024)
What's documented in AI? Systematic Analysis of 32K AI Model Cards
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
What Makes Code Generation Ethically Sourced?
by: Xu, Zhuolin, et al.
Published: (2025)
by: Xu, Zhuolin, et al.
Published: (2025)
What can Large Language Models Capture about Code Functional Equivalence?
by: Maveli, Nickil, et al.
Published: (2024)
by: Maveli, Nickil, et al.
Published: (2024)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
Are LLMs Reliable Code Reviewers? Systematic Overcorrection in Requirement Conformance Judgement
by: Jin, Haolin, et al.
Published: (2026)
by: Jin, Haolin, et al.
Published: (2026)
CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model
by: Di, Peng, et al.
Published: (2023)
by: Di, Peng, et al.
Published: (2023)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
by: Wu, Xingfang, et al.
Published: (2024)
by: Wu, Xingfang, et al.
Published: (2024)
Operational Robustness of LLMs on Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2026)
by: Paul, Debalina Ghosh, et al.
Published: (2026)
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration
by: Shen, Yuanhao, et al.
Published: (2024)
by: Shen, Yuanhao, et al.
Published: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
by: Xu, Weiwei, et al.
Published: (2024)
by: Xu, Weiwei, et al.
Published: (2024)
Is ChatGPT a Good Software Librarian? An Exploratory Study on the Use of ChatGPT for Software Library Recommendations
by: Latendresse, Jasmine, et al.
Published: (2024)
by: Latendresse, Jasmine, et al.
Published: (2024)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
Modeling Code: Is Text All You Need?
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis
by: Gajjar, Jugal
Published: (2026)
by: Gajjar, Jugal
Published: (2026)
What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
by: Robredo, Mikel, et al.
Published: (2025)
by: Robredo, Mikel, et al.
Published: (2025)
Distilled GPT for Source Code Summarization
by: Su, Chia-Yi, et al.
Published: (2023)
by: Su, Chia-Yi, et al.
Published: (2023)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
by: Jiang, Shan, et al.
Published: (2026)
by: Jiang, Shan, et al.
Published: (2026)
ChatGPT Incorrectness Detection in Software Reviews
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
by: Bian, Yutong, et al.
Published: (2025)
by: Bian, Yutong, et al.
Published: (2025)
Can ChatGPT Support Developers? An Empirical Evaluation of Large Language Models for Code Generation
by: Jin, Kailun, et al.
Published: (2024)
by: Jin, Kailun, et al.
Published: (2024)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
What You Use is What You Get: Unforced Errors in Studying Cultural Aspects in Agile Software Development
by: Neumann, Michael, et al.
Published: (2024)
by: Neumann, Michael, et al.
Published: (2024)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
Code Generation by Differential Test Time Scaling
by: He, Yifeng, et al.
Published: (2026)
by: He, Yifeng, et al.
Published: (2026)
The Counterfeit Conundrum: Can Code Language Models Grasp the Nuances of Their Incorrect Generations?
by: Gu, Alex, et al.
Published: (2024)
by: Gu, Alex, et al.
Published: (2024)
Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Similar Items
-
Your Code Agent Can Grow Alongside You with Structured Memory
by: Deng, Yi-Xuan, et al.
Published: (2026) -
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
by: Spracklen, Joseph, et al.
Published: (2024) -
Let the Code LLM Edit Itself When You Edit the Code
by: He, Zhenyu, et al.
Published: (2024) -
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
by: Huang, Linghan, et al.
Published: (2025) -
What You See Is What You Get: Attention-based Self-guided Automatic Unit Test Generation
by: Yin, Xin, et al.
Published: (2024)