Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gong, Linyuan, Wang, Sida, Elhoushi, Mostafa, Cheung, Alvin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Structure-Aware Fill-in-the-Middle Pretraining for Code
par: Gong, Linyuan, et autres
Publié: (2025)
par: Gong, Linyuan, et autres
Publié: (2025)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
par: Gong, Linyuan, et autres
Publié: (2024)
par: Gong, Linyuan, et autres
Publié: (2024)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
par: Skurla, Adam, et autres
Publié: (2026)
par: Skurla, Adam, et autres
Publié: (2026)
How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data
par: Wang, Yejie, et autres
Publié: (2024)
par: Wang, Yejie, et autres
Publié: (2024)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
par: Yang, Dayu, et autres
Publié: (2025)
par: Yang, Dayu, et autres
Publié: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
par: Dong, Yihong, et autres
Publié: (2025)
par: Dong, Yihong, et autres
Publié: (2025)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
par: Wang, Xinchen, et autres
Publié: (2025)
par: Wang, Xinchen, et autres
Publié: (2025)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
par: Guo, Jiawei, et autres
Publié: (2024)
par: Guo, Jiawei, et autres
Publié: (2024)
Vibe Checker: Aligning Code Evaluation with Human Preference
par: Zhong, Ming, et autres
Publié: (2025)
par: Zhong, Ming, et autres
Publié: (2025)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
par: Lu, Yifei, et autres
Publié: (2025)
par: Lu, Yifei, et autres
Publié: (2025)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
par: Zheng, Zihan, et autres
Publié: (2025)
par: Zheng, Zihan, et autres
Publié: (2025)
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
par: Tantithamthavorn, Kla, et autres
Publié: (2026)
par: Tantithamthavorn, Kla, et autres
Publié: (2026)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
par: Dai, Hankun, et autres
Publié: (2025)
par: Dai, Hankun, et autres
Publié: (2025)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
par: Ficek, Aleksander, et autres
Publié: (2025)
par: Ficek, Aleksander, et autres
Publié: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
par: Shetty, Manish, et autres
Publié: (2025)
par: Shetty, Manish, et autres
Publié: (2025)
DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
par: Lu, Junyi, et autres
Publié: (2024)
par: Lu, Junyi, et autres
Publié: (2024)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
par: Cao, Jialun, et autres
Publié: (2024)
par: Cao, Jialun, et autres
Publié: (2024)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
par: Wei, Yuxiang, et autres
Publié: (2025)
par: Wei, Yuxiang, et autres
Publié: (2025)
CODEMENV: Benchmarking Large Language Models on Code Migration
par: Cheng, Keyuan, et autres
Publié: (2025)
par: Cheng, Keyuan, et autres
Publié: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
par: Fan, Lishui, et autres
Publié: (2025)
par: Fan, Lishui, et autres
Publié: (2025)
Let the Code LLM Edit Itself When You Edit the Code
par: He, Zhenyu, et autres
Publié: (2024)
par: He, Zhenyu, et autres
Publié: (2024)
From I/O to Code with Discovery Agent
par: Dong, Yihong, et autres
Publié: (2026)
par: Dong, Yihong, et autres
Publié: (2026)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
par: Yu, Zhuohao, et autres
Publié: (2024)
par: Yu, Zhuohao, et autres
Publié: (2024)
A Survey on Code Generation with LLM-based Agents
par: Dong, Yihong, et autres
Publié: (2025)
par: Dong, Yihong, et autres
Publié: (2025)
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
par: Liu, Sicong, et autres
Publié: (2026)
par: Liu, Sicong, et autres
Publié: (2026)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
par: Hu, Ruida, et autres
Publié: (2025)
par: Hu, Ruida, et autres
Publié: (2025)
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
par: Xu, Derek, et autres
Publié: (2024)
par: Xu, Derek, et autres
Publié: (2024)
Selective Prompt Anchoring for Code Generation
par: Tian, Yuan, et autres
Publié: (2024)
par: Tian, Yuan, et autres
Publié: (2024)
MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
par: Kadosh, Tal, et autres
Publié: (2023)
par: Kadosh, Tal, et autres
Publié: (2023)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
par: Wong, Sherman, et autres
Publié: (2025)
par: Wong, Sherman, et autres
Publié: (2025)
Scaling Test-Time Compute for Agentic Coding
par: Kim, Joongwon, et autres
Publié: (2026)
par: Kim, Joongwon, et autres
Publié: (2026)
AuPair: Golden Example Pairs for Code Repair
par: Mavalankar, Aditi, et autres
Publié: (2025)
par: Mavalankar, Aditi, et autres
Publié: (2025)
Towards Practical Defect-Focused Automated Code Review
par: Lu, Junyi, et autres
Publié: (2025)
par: Lu, Junyi, et autres
Publié: (2025)
code_transformed: The Influence of Large Language Models on Code
par: Xu, Yuliang, et autres
Publié: (2025)
par: Xu, Yuliang, et autres
Publié: (2025)
Improving Code Generation by Training with Natural Language Feedback
par: Chen, Angelica, et autres
Publié: (2023)
par: Chen, Angelica, et autres
Publié: (2023)
Investigating the Efficacy of Large Language Models for Code Clone Detection
par: Khajezade, Mohamad, et autres
Publié: (2024)
par: Khajezade, Mohamad, et autres
Publié: (2024)
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
par: Wu, Yangzhen, et autres
Publié: (2026)
par: Wu, Yangzhen, et autres
Publié: (2026)
SQLong: Enhanced NL2SQL for Longer Contexts with LLMs
par: Nguyen, Dai Quoc, et autres
Publié: (2025)
par: Nguyen, Dai Quoc, et autres
Publié: (2025)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
par: Hassid, Michael, et autres
Publié: (2024)
par: Hassid, Michael, et autres
Publié: (2024)
Documents similaires
-
Structure-Aware Fill-in-the-Middle Pretraining for Code
par: Gong, Linyuan, et autres
Publié: (2025) -
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
par: Gong, Linyuan, et autres
Publié: (2024) -
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
par: Skurla, Adam, et autres
Publié: (2026) -
How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data
par: Wang, Yejie, et autres
Publié: (2024) -
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
par: Yang, Dayu, et autres
Publié: (2025)