Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Allamanis, Miltiadis, Panthaplackel, Sheena, Yin, Pengcheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disproving Program Equivalence with LLMs
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
von: Ruiz, Fernando Vallecillos, et al.
Veröffentlicht: (2024)
von: Ruiz, Fernando Vallecillos, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
On LLMs' Internal Representation of Code Correctness
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
Calibration and Correctness of Language Models for Code
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
Teaching Code Refactoring Using LLMs
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
Towards Verified Code Reasoning by LLMs
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
von: Mhatre, Akshay, et al.
Veröffentlicht: (2025)
von: Mhatre, Akshay, et al.
Veröffentlicht: (2025)
Evaluating the Use of LLMs for Documentation to Code Traceability
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
Grounding Data Science Code Generation with Input-Output Specifications
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
Ensuring Functional Correctness of Large Code Models with Selective Generation
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024)
von: Li, Linyi, et al.
Veröffentlicht: (2024)
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2026)
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2026)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications
von: Huynh, Nam, et al.
Veröffentlicht: (2025)
von: Huynh, Nam, et al.
Veröffentlicht: (2025)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
Correctness-Guaranteed Code Generation via Constrained Decoding
von: Li, Lingxiao, et al.
Veröffentlicht: (2025)
von: Li, Lingxiao, et al.
Veröffentlicht: (2025)
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems
von: Xing, Jun, et al.
Veröffentlicht: (2025)
von: Xing, Jun, et al.
Veröffentlicht: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Disproving Program Equivalence with LLMs
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025) -
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024) -
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
von: Hooda, Ashish, et al.
Veröffentlicht: (2024) -
Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
von: Ruiz, Fernando Vallecillos, et al.
Veröffentlicht: (2024) -
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)