CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Naik, Atharva, Alenius, Marcus, Fried, Daniel, Rose, Carolyn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
von: Kapadnis, Manav Nitin, et al.
Veröffentlicht: (2025)
von: Kapadnis, Manav Nitin, et al.
Veröffentlicht: (2025)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
von: Naik, Atharva
Veröffentlicht: (2024)
von: Naik, Atharva
Veröffentlicht: (2024)
Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding
von: Di, Yifeng, et al.
Veröffentlicht: (2025)
von: Di, Yifeng, et al.
Veröffentlicht: (2025)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
von: Sehgal, Atharva, et al.
Veröffentlicht: (2026)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2026)
DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
AI-Assisted Fixes to Code Review Comments at Scale
von: Maddila, Chandra, et al.
Veröffentlicht: (2025)
von: Maddila, Chandra, et al.
Veröffentlicht: (2025)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
von: Pereira, Kristen, et al.
Veröffentlicht: (2026)
von: Pereira, Kristen, et al.
Veröffentlicht: (2026)
Code Broker: A Multi-Agent System for Automated Code Quality Assessment
von: Attrah, Samer
Veröffentlicht: (2026)
von: Attrah, Samer
Veröffentlicht: (2026)
Towards Practical Defect-Focused Automated Code Review
von: Lu, Junyi, et al.
Veröffentlicht: (2025)
von: Lu, Junyi, et al.
Veröffentlicht: (2025)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
Semantically Aligned Question and Code Generation for Automated Insight Generation
von: Singha, Ananya, et al.
Veröffentlicht: (2024)
von: Singha, Ananya, et al.
Veröffentlicht: (2024)
SEW: Self-Evolving Agentic Workflows for Automated Code Generation
von: Liu, Siwei, et al.
Veröffentlicht: (2025)
von: Liu, Siwei, et al.
Veröffentlicht: (2025)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
von: Kammakomati, Mehant, et al.
Veröffentlicht: (2024)
von: Kammakomati, Mehant, et al.
Veröffentlicht: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Chain of Grounded Objectives: Bridging Process and Goal-oriented Prompting for Code Generation
von: Yeo, Sangyeop, et al.
Veröffentlicht: (2025)
von: Yeo, Sangyeop, et al.
Veröffentlicht: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
Code Sharing In Prediction Model Research: A Scoping Review
von: Sounack, Thomas, et al.
Veröffentlicht: (2026)
von: Sounack, Thomas, et al.
Veröffentlicht: (2026)
AI-Mediated Code Comment Improvement
von: Dhakal, Maria, et al.
Veröffentlicht: (2025)
von: Dhakal, Maria, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
Specification and Detection of LLM Code Smells
von: Mahmoudi, Brahim, et al.
Veröffentlicht: (2025)
von: Mahmoudi, Brahim, et al.
Veröffentlicht: (2025)
Investigating The Smells of LLM Generated Code
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2025)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2025)
Rethinking Code Refinement: Learning to Judge Code Efficiency
von: Seo, Minju, et al.
Veröffentlicht: (2024)
von: Seo, Minju, et al.
Veröffentlicht: (2024)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
DevEval: Evaluating Code Generation in Practical Software Projects
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
ICE-Score: Instructing Large Language Models to Evaluate Code
von: Zhuo, Terry Yue
Veröffentlicht: (2023)
von: Zhuo, Terry Yue
Veröffentlicht: (2023)
McMining: Automated Discovery of Misconceptions in Student Code
von: Al-Hossami, Erfan, et al.
Veröffentlicht: (2025)
von: Al-Hossami, Erfan, et al.
Veröffentlicht: (2025)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
von: Kapadnis, Manav Nitin, et al.
Veröffentlicht: (2025) -
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023) -
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025) -
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025) -
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
von: Naik, Atharva
Veröffentlicht: (2024)