BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuo, Terry Yue, Vu, Minh Chien, Chim, Jenny, Hu, Han, Yu, Wenhao, Widyasari, Ratnadira, Yusuf, Imam Nur Bani, Zhan, Haolan, He, Junda, Paul, Indraneil, Brunner, Simon, Gong, Chen, Hoang, Thong, Zebaze, Armel Randy, Hong, Xiaoheng, Li, Wen-Ding, Kaddour, Jean, Xu, Ming, Zhang, Zhihan, Yadav, Prateek, Jain, Naman, Gu, Alex, Cheng, Zhoujun, Liu, Jiawei, Liu, Qian, Wang, Zijian, Hui, Binyuan, Muennighoff, Niklas, Lo, David, Fried, Daniel, Du, Xiaoning, de Vries, Harm, Von Werra, Leandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
OctoPack: Instruction Tuning Code Large Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
Explaining Explanation: An Empirical Study on Explanation in Code Reviews
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2023)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2023)
Beyond ChatGPT: Enhancing Software Quality Assurance Tasks with Diverse LLMs and Validation Techniques
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
Demystifying Faulty Code with LLM: Step-by-Step Reasoning for Explainable Fault Localization
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
SelfCodeAlign: Self-Alignment for Code Generation
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
von: Liu, Yue, et al.
Veröffentlicht: (2026)
von: Liu, Yue, et al.
Veröffentlicht: (2026)
Transducer Tuning: Efficient Model Adaptation for Software Tasks Using Code Property Graphs
von: Yusuf, Imam Nur Bani, et al.
Veröffentlicht: (2024)
von: Yusuf, Imam Nur Bani, et al.
Veröffentlicht: (2024)
Tree of Problems: Improving structured problem solving with compositionality
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Aletheia: What Makes RLVR For Code Verifiers Tick?
von: Venkatkrishna, Vatsal, et al.
Veröffentlicht: (2026)
von: Venkatkrishna, Vatsal, et al.
Veröffentlicht: (2026)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
von: Paul, Indraneil, et al.
Veröffentlicht: (2026)
von: Paul, Indraneil, et al.
Veröffentlicht: (2026)
Turning the Tide: Repository-based Code Reflection
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Code Review Agent Benchmark
von: Zhang, Yuntong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuntong, et al.
Veröffentlicht: (2026)
IFEvalCode: Controlled Code Generation
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel
von: Lyu, Yunbo, et al.
Veröffentlicht: (2023)
von: Lyu, Yunbo, et al.
Veröffentlicht: (2023)
Disentangling meaning from language in LLM-based machine translation
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
Compiling Code LLMs into Lightweight Executables
von: Shi, Jieke, et al.
Veröffentlicht: (2026)
von: Shi, Jieke, et al.
Veröffentlicht: (2026)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
Your Instructions Are Not Always Helpful: Assessing the Efficacy of Instruction Fine-tuning for Software Vulnerability Detection
von: Yusuf, Imam Nur Bani, et al.
Veröffentlicht: (2024)
von: Yusuf, Imam Nur Bani, et al.
Veröffentlicht: (2024)
Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
von: Danassis, Panayiotis, et al.
Veröffentlicht: (2025)
von: Danassis, Panayiotis, et al.
Veröffentlicht: (2025)
From Code to Courtroom: LLMs as the New Software Judges
von: He, Junda, et al.
Veröffentlicht: (2025)
von: He, Junda, et al.
Veröffentlicht: (2025)
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
von: Lyu, Yunbo, et al.
Veröffentlicht: (2026)
von: Lyu, Yunbo, et al.
Veröffentlicht: (2026)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
Towards Better Correctness and Efficiency in Code Generation
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
Active Sensing with Predictive Coding and Uncertainty Minimization
von: Sharafeldin, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Sharafeldin, Abdelrahman, et al.
Veröffentlicht: (2023)
Target Policy Optimization
von: Kaddour, Jean
Veröffentlicht: (2026)
von: Kaddour, Jean
Veröffentlicht: (2026)
La participación ciudadana en la planificación local y urbana en Argelia
von: Kaddour Derbal
Veröffentlicht: (2022)
von: Kaddour Derbal
Veröffentlicht: (2022)
Aligned Multi-View Scripts for Universal Chart-to-Code Generation
von: Zhang, Zhihan, et al.
Veröffentlicht: (2026)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2026)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
Evaluating and Achieving Controllable Code Completion in Code LLM
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
Lemur: Harmonizing Natural Language and Code for Language Agents
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024) -
OctoPack: Instruction Tuning Code Large Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023) -
Explaining Explanation: An Empirical Study on Explanation in Code Reviews
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2023) -
Beyond ChatGPT: Enhancing Software Quality Assurance Tasks with Diverse LLMs and Validation Techniques
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024) -
Demystifying Faulty Code with LLM: Step-by-Step Reasoning for Explainable Fault Localization
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)