Planning-Aware Code Infilling via Horizon-Length Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Yifeng, Ding, Hantian, Wang, Shiqi, Sun, Qing, Kumar, Varun, Wang, Zijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Code Preference via Synthetic Evolution
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
Evaluating Language Models for Efficient Code Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
RepoQA: Evaluating Long Context Code Understanding
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
SelfCodeAlign: Self-Alignment for Code Generation
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
von: Li, Yufei, et al.
Veröffentlicht: (2024)
von: Li, Yufei, et al.
Veröffentlicht: (2024)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
von: Nathan, Varun, et al.
Veröffentlicht: (2026)
von: Nathan, Varun, et al.
Veröffentlicht: (2026)
Suggesting Code Edits in Interactive Machine Learning Notebooks Using Large Language Models
von: Jin, Bihui, et al.
Veröffentlicht: (2025)
von: Jin, Bihui, et al.
Veröffentlicht: (2025)
Towards Understanding What Code Language Models Learned
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
Position: Intelligent Coding Systems Should Write Programs with Justifications
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2025)
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2025)
IntentCoding: Amplifying User Intent in Code Generation
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
von: Lu, Junyi, et al.
Veröffentlicht: (2024)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
von: Park, Chansung, et al.
Veröffentlicht: (2026)
von: Park, Chansung, et al.
Veröffentlicht: (2026)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
von: Weyssow, Martin, et al.
Veröffentlicht: (2024)
von: Weyssow, Martin, et al.
Veröffentlicht: (2024)
HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench
von: Wang, Yueyang, et al.
Veröffentlicht: (2026)
von: Wang, Yueyang, et al.
Veröffentlicht: (2026)
Revisiting the Plastic Surgery Hypothesis via Large Language Models
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2023)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2023)
Magicoder: Empowering Code Generation with OSS-Instruct
von: Wei, Yuxiang, et al.
Veröffentlicht: (2023)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2023)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
Wisdom and Delusion of LLM Ensembles for Code Generation and Repair
von: Vallecillos-Ruiz, Fernando, et al.
Veröffentlicht: (2025)
von: Vallecillos-Ruiz, Fernando, et al.
Veröffentlicht: (2025)
GiFT: Gibbs Fine-Tuning for Code Generation
von: Li, Haochen, et al.
Veröffentlicht: (2025)
von: Li, Haochen, et al.
Veröffentlicht: (2025)
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Code Preference via Synthetic Evolution
von: Liu, Jiawei, et al.
Veröffentlicht: (2024) -
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025) -
Evaluating Language Models for Efficient Code Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2024) -
RepoQA: Evaluating Long Context Code Understanding
von: Liu, Jiawei, et al.
Veröffentlicht: (2024) -
SelfCodeAlign: Self-Alignment for Code Generation
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)