Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
Fuente:
arXiv
Saved in:
| Main Authors: | He, Fusen, Zhai, Juan, Pan, Minxue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying and Assessing Code Understandability in Java Decompilation
by: Qin, Ruixin, et al.
Published: (2024)
by: Qin, Ruixin, et al.
Published: (2024)
LLMCup: Ranking-Enhanced Comment Updating with LLMs
by: Ge, Hua, et al.
Published: (2025)
by: Ge, Hua, et al.
Published: (2025)
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
by: Zhang, Gehao, et al.
Published: (2026)
by: Zhang, Gehao, et al.
Published: (2026)
Breaking the Myth: Can Small Models Infer Postconditions Too?
by: Zhang, Gehao, et al.
Published: (2025)
by: Zhang, Gehao, et al.
Published: (2025)
Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation
by: Sharma, Arindam, et al.
Published: (2025)
by: Sharma, Arindam, et al.
Published: (2025)
CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation
by: Pan, Ruwei, et al.
Published: (2025)
by: Pan, Ruwei, et al.
Published: (2025)
Assessing Code Generation with Intermediate Languages
by: Deng, Xun, et al.
Published: (2024)
by: Deng, Xun, et al.
Published: (2024)
Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation
by: Sarker, Laboni, et al.
Published: (2024)
by: Sarker, Laboni, et al.
Published: (2024)
Assessing the Impact of Requirement Ambiguity on LLM-based Function-Level Code Generation
by: Yang, Di, et al.
Published: (2026)
by: Yang, Di, et al.
Published: (2026)
Inducing Vulnerable Code Generation in LLM Coding Assistants
by: Zeng, Binqi, et al.
Published: (2025)
by: Zeng, Binqi, et al.
Published: (2025)
Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness
by: Blyth, Scott, et al.
Published: (2025)
by: Blyth, Scott, et al.
Published: (2025)
Beyond Code Similarity: Benchmarking the Plausibility, Efficiency, and Complexity of LLM-Generated Smart Contracts
by: Salzano, Francesco, et al.
Published: (2025)
by: Salzano, Francesco, et al.
Published: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
Using Semantic Distance to Estimate Uncertainty in LLM-Based Code Generation
by: He, Weilin, et al.
Published: (2026)
by: He, Weilin, et al.
Published: (2026)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
When is Generated Code Difficult to Comprehend? Assessing AI Agent Python Code Proficiency in the Wild
by: Temkulkiat, Nanthit, et al.
Published: (2026)
by: Temkulkiat, Nanthit, et al.
Published: (2026)
Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond
by: Le-Anh, Minh, et al.
Published: (2026)
by: Le-Anh, Minh, et al.
Published: (2026)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
by: Jia, Haoxiang, et al.
Published: (2025)
by: Jia, Haoxiang, et al.
Published: (2025)
Beyond Strict Rules: Assessing the Effectiveness of Large Language Models for Code Smell Detection
by: Souza, Saymon, et al.
Published: (2026)
by: Souza, Saymon, et al.
Published: (2026)
Unsafe and Unused? A History of Utility Code in Mature Open Source Projects
by: Keller, Brandon, et al.
Published: (2026)
by: Keller, Brandon, et al.
Published: (2026)
Assessing AI-Based Code Assistants in Method Generation Tasks
by: Corso, Vincenzo, et al.
Published: (2024)
by: Corso, Vincenzo, et al.
Published: (2024)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
by: Manh, Dung Nguyen, et al.
Published: (2024)
by: Manh, Dung Nguyen, et al.
Published: (2024)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
by: Hasan, Md Mahade, et al.
Published: (2025)
by: Hasan, Md Mahade, et al.
Published: (2025)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
by: He, Yicheng, et al.
Published: (2026)
by: He, Yicheng, et al.
Published: (2026)
Modularization is Better: Effective Code Generation with Modular Prompting
by: Pan, Ruwei, et al.
Published: (2025)
by: Pan, Ruwei, et al.
Published: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)
by: Chen, Valerie, et al.
Published: (2026)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
by: Crupi, Giuseppe, et al.
Published: (2025)
by: Crupi, Giuseppe, et al.
Published: (2025)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
by: Zhang, Yuanliang, et al.
Published: (2024)
by: Zhang, Yuanliang, et al.
Published: (2024)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
CodeScore: Evaluating Code Generation by Learning Code Execution
by: Dong, Yihong, et al.
Published: (2023)
by: Dong, Yihong, et al.
Published: (2023)
Contextualized Code Pretraining for Code Generation
by: Liu, Chen, et al.
Published: (2026)
by: Liu, Chen, et al.
Published: (2026)
Context-Aware CodeLLM Eviction for AI-assisted Coding
by: Thangarajah, Kishanthan, et al.
Published: (2025)
by: Thangarajah, Kishanthan, et al.
Published: (2025)
Similar Items
-
Demystifying and Assessing Code Understandability in Java Decompilation
by: Qin, Ruixin, et al.
Published: (2024) -
LLMCup: Ranking-Enhanced Comment Updating with LLMs
by: Ge, Hua, et al.
Published: (2025) -
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
by: Zhang, Gehao, et al.
Published: (2026) -
Breaking the Myth: Can Small Models Infer Postconditions Too?
by: Zhang, Gehao, et al.
Published: (2025) -
Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation
by: Sharma, Arindam, et al.
Published: (2025)