CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Kaiwen, Guo, Hongcheng, Shi, Xuanqing, Cao, Shaosheng, Di, Donglin, Li, Zhoujun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
von: Yang, Rem, et al.
Veröffentlicht: (2025)
von: Yang, Rem, et al.
Veröffentlicht: (2025)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024)
von: Li, Linyi, et al.
Veröffentlicht: (2024)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
CODEMENV: Benchmarking Large Language Models on Code Migration
von: Cheng, Keyuan, et al.
Veröffentlicht: (2025)
von: Cheng, Keyuan, et al.
Veröffentlicht: (2025)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
von: Islah, Nizar, et al.
Veröffentlicht: (2024)
von: Islah, Nizar, et al.
Veröffentlicht: (2024)
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
von: Cassano, Federico, et al.
Veröffentlicht: (2023)
von: Cassano, Federico, et al.
Veröffentlicht: (2023)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language Models
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
von: Park, Chansung, et al.
Veröffentlicht: (2026)
von: Park, Chansung, et al.
Veröffentlicht: (2026)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
von: Li, Yufei, et al.
Veröffentlicht: (2024)
von: Li, Yufei, et al.
Veröffentlicht: (2024)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications
von: Huynh, Nam, et al.
Veröffentlicht: (2025)
von: Huynh, Nam, et al.
Veröffentlicht: (2025)
Ensuring Functional Correctness of Large Code Models with Selective Generation
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model
von: Di, Peng, et al.
Veröffentlicht: (2023)
von: Di, Peng, et al.
Veröffentlicht: (2023)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2024)
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2024)
Calibration and Correctness of Language Models for Code
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
MLAD: A Unified Model for Multi-system Log Anomaly Detection
von: Zang, Runqiang, et al.
Veröffentlicht: (2024)
von: Zang, Runqiang, et al.
Veröffentlicht: (2024)
Unlearning Trojans in Large Language Models: A Comparison Between Natural Language and Source Code
von: Kazemi, Mahdi, et al.
Veröffentlicht: (2024)
von: Kazemi, Mahdi, et al.
Veröffentlicht: (2024)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
von: Jiang, Mingchao, et al.
Veröffentlicht: (2025)
von: Jiang, Mingchao, et al.
Veröffentlicht: (2025)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
On Trojan Signatures in Large Language Models of Code
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
Follow Your Nose -- Which Code Smells are Worth Chasing?
von: Amit, Idan, et al.
Veröffentlicht: (2021)
von: Amit, Idan, et al.
Veröffentlicht: (2021)
Applying the Chinese Wall Reverse Engineering Technique to Large Language Model Code Editing
von: Hanmongkolchai, Manatsawin
Veröffentlicht: (2025)
von: Hanmongkolchai, Manatsawin
Veröffentlicht: (2025)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
SemRep: Generative Code Representation Learning with Code Transformations
von: Li, Weichen, et al.
Veröffentlicht: (2026)
von: Li, Weichen, et al.
Veröffentlicht: (2026)
LLaMoCo: Instruction Tuning of Large Language Models for Optimization Code Generation
von: Ma, Zeyuan, et al.
Veröffentlicht: (2024)
von: Ma, Zeyuan, et al.
Veröffentlicht: (2024)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
von: Weyssow, Martin, et al.
Veröffentlicht: (2023)
von: Weyssow, Martin, et al.
Veröffentlicht: (2023)
Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and Analysis
von: Gai, Jiahao, et al.
Veröffentlicht: (2025)
von: Gai, Jiahao, et al.
Veröffentlicht: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025) -
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
von: Yang, Rem, et al.
Veröffentlicht: (2025) -
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024) -
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024) -
CODEMENV: Benchmarking Large Language Models on Code Migration
von: Cheng, Keyuan, et al.
Veröffentlicht: (2025)