Inferring Code Correctness from Specification
Fuente:
arXiv
Saved in:
| Main Authors: | Florian, Tambon, Mike, Papadakis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
by: He, Yicheng, et al.
Published: (2026)
by: He, Yicheng, et al.
Published: (2026)
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
by: Dong, Zeming, et al.
Published: (2024)
by: Dong, Zeming, et al.
Published: (2024)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
by: Dong, Zeming, et al.
Published: (2023)
by: Dong, Zeming, et al.
Published: (2023)
Specification and Detection of LLM Code Smells
by: Mahmoudi, Brahim, et al.
Published: (2025)
by: Mahmoudi, Brahim, et al.
Published: (2025)
ML Code Smells: From Specification to Detection
by: Mahmoudi, Brahim, et al.
Published: (2025)
by: Mahmoudi, Brahim, et al.
Published: (2025)
Learning Generalizable Multimodal Representations for Software Vulnerability Detection
by: Dong, Zeming, et al.
Published: (2026)
by: Dong, Zeming, et al.
Published: (2026)
Automatic Identification of Machine Learning-Specific Code Smells
by: Hamfelt, Peter, et al.
Published: (2025)
by: Hamfelt, Peter, et al.
Published: (2025)
Unveiling Project-Specific Bias in Neural Code Models
by: Li, Zhiming, et al.
Published: (2022)
by: Li, Zhiming, et al.
Published: (2022)
BabelCoder: Agentic Code Translation with Specification Alignment
by: Rabbi, Fazle, et al.
Published: (2025)
by: Rabbi, Fazle, et al.
Published: (2025)
Towards Better Correctness and Efficiency in Code Generation
by: Feng, Yunlong, et al.
Published: (2025)
by: Feng, Yunlong, et al.
Published: (2025)
Can Code Evaluation Metrics Detect Code Plagiarism?
by: Ebrahim, Fahad, et al.
Published: (2026)
by: Ebrahim, Fahad, et al.
Published: (2026)
Inferring Pluggable Types with Machine Learning
by: Siddiqui, Kazi Amanul Islam, et al.
Published: (2024)
by: Siddiqui, Kazi Amanul Islam, et al.
Published: (2024)
Benchmarking Correctness and Security in Multi-Turn Code Generation
by: Rawal, Ruchit, et al.
Published: (2025)
by: Rawal, Ruchit, et al.
Published: (2025)
VeriAct: Beyond Verifiability -- Agentic Synthesis of Correct and Complete Formal Specifications
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
by: Yan, Lu, et al.
Published: (2026)
by: Yan, Lu, et al.
Published: (2026)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
by: Kuhar, Sachit, et al.
Published: (2024)
by: Kuhar, Sachit, et al.
Published: (2024)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
Automating the Correctness Assessment of AI-generated Code for Security Contexts
by: Cotroneo, Domenico, et al.
Published: (2023)
by: Cotroneo, Domenico, et al.
Published: (2023)
Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code
by: Rajput, Prateek, et al.
Published: (2026)
by: Rajput, Prateek, et al.
Published: (2026)
HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
by: Naik, Atharva
Published: (2024)
by: Naik, Atharva
Published: (2024)
DomAgent: Leveraging Knowledge Graphs and Case-Based Reasoning for Domain-Specific Code Generation
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
by: Pathak, Aditya, et al.
Published: (2025)
by: Pathak, Aditya, et al.
Published: (2025)
Breaking the Myth: Can Small Models Infer Postconditions Too?
by: Zhang, Gehao, et al.
Published: (2025)
by: Zhang, Gehao, et al.
Published: (2025)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
by: Vartziotis, Tina, et al.
Published: (2024)
by: Vartziotis, Tina, et al.
Published: (2024)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
by: Khati, Dipin, et al.
Published: (2026)
by: Khati, Dipin, et al.
Published: (2026)
Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
MergeRepair: An Exploratory Study on Merging Task-Specific Adapters in Code LLMs for Automated Program Repair
by: Dehghan, Meghdad, et al.
Published: (2024)
by: Dehghan, Meghdad, et al.
Published: (2024)
Enhancing Cross-Language Code Translation via Task-Specific Embedding Alignment in Retrieval-Augmented Generation
by: Bhattarai, Manish, et al.
Published: (2024)
by: Bhattarai, Manish, et al.
Published: (2024)
LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness
by: Ye, Tong, et al.
Published: (2025)
by: Ye, Tong, et al.
Published: (2025)
AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code
by: Fathallah, Nadeen, et al.
Published: (2025)
by: Fathallah, Nadeen, et al.
Published: (2025)
Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions
by: Wang, Chengjie, et al.
Published: (2026)
by: Wang, Chengjie, et al.
Published: (2026)
Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
by: Ashiga, Mari, et al.
Published: (2025)
by: Ashiga, Mari, et al.
Published: (2025)
Similar Items
-
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024) -
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024) -
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
by: Akli, Amal, et al.
Published: (2026) -
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
by: He, Yicheng, et al.
Published: (2026)