CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Jinjun, Cui, Leyi, Huang, Kele, Yang, Junfeng, Ray, Baishakhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
von: Peng, Jinjun, et al.
Veröffentlicht: (2026)
von: Peng, Jinjun, et al.
Veröffentlicht: (2026)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
A Multi-Perspective Architecture for Semantic Code Search
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2020)
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2020)
CYCLE: Learning to Self-Refine the Code Generation
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
EditLord: Learning Code Transformation Rules for Code Editing
von: Li, Weichen, et al.
Veröffentlicht: (2025)
von: Li, Weichen, et al.
Veröffentlicht: (2025)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
von: Yang, Rem, et al.
Veröffentlicht: (2025)
von: Yang, Rem, et al.
Veröffentlicht: (2025)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
Wisdom and Delusion of LLM Ensembles for Code Generation and Repair
von: Vallecillos-Ruiz, Fernando, et al.
Veröffentlicht: (2025)
von: Vallecillos-Ruiz, Fernando, et al.
Veröffentlicht: (2025)
Evaluating Language Models for Efficient Code Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
von: Min, Marcus J., et al.
Veröffentlicht: (2023)
von: Min, Marcus J., et al.
Veröffentlicht: (2023)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
von: Ryan, Gabriel, et al.
Veröffentlicht: (2024)
von: Ryan, Gabriel, et al.
Veröffentlicht: (2024)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
von: Park, Chansung, et al.
Veröffentlicht: (2026)
von: Park, Chansung, et al.
Veröffentlicht: (2026)
Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis
von: Murphy, William, et al.
Veröffentlicht: (2024)
von: Murphy, William, et al.
Veröffentlicht: (2024)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
RepoQA: Evaluating Long Context Code Understanding
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
IntentCoding: Amplifying User Intent in Code Generation
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
SelfCodeAlign: Self-Alignment for Code Generation
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
von: Weyssow, Martin, et al.
Veröffentlicht: (2024)
von: Weyssow, Martin, et al.
Veröffentlicht: (2024)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
A Survey on Code Generation with LLM-based Agents
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
GiFT: Gibbs Fine-Tuning for Code Generation
von: Li, Haochen, et al.
Veröffentlicht: (2025)
von: Li, Haochen, et al.
Veröffentlicht: (2025)
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
Test Code Generation for Telecom Software Systems using Two-Stage Generative Model
von: Nabeel, Mohamad, et al.
Veröffentlicht: (2024)
von: Nabeel, Mohamad, et al.
Veröffentlicht: (2024)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
The Impact of Prompt Programming on Function-Level Code Generation
von: Khojah, Ranim, et al.
Veröffentlicht: (2024)
von: Khojah, Ranim, et al.
Veröffentlicht: (2024)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
Code Review Without Borders: Evaluating Synthetic vs. Real Data for Review Recommendation
von: Cohen, Yogev, et al.
Veröffentlicht: (2025)
von: Cohen, Yogev, et al.
Veröffentlicht: (2025)
Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
von: Ridnik, Tal, et al.
Veröffentlicht: (2024)
von: Ridnik, Tal, et al.
Veröffentlicht: (2024)
SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
von: Peng, Jinjun, et al.
Veröffentlicht: (2026) -
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024) -
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
von: Chen, Simin, et al.
Veröffentlicht: (2025) -
A Multi-Perspective Architecture for Semantic Code Search
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2020) -
CYCLE: Learning to Self-Refine the Code Generation
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)