Verification Limits Code LLM Training
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gureja, Srishti, Tommasone, Elena, He, Jingyi, Hooker, Sara, Gallé, Matthias, Fadaee, Marzieh |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Pull Requests as a Training Signal for Repo-Level Code Editing
par: Zhu, Qinglin, et autres
Publié: (2026)
par: Zhu, Qinglin, et autres
Publié: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
Training Versatile Coding Agents in Synthetic Environments
par: Zhu, Yiqi, et autres
Publié: (2025)
par: Zhu, Yiqi, et autres
Publié: (2025)
Ranking LLM-Generated Loop Invariants for Program Verification
par: Chakraborty, Saikat, et autres
Publié: (2023)
par: Chakraborty, Saikat, et autres
Publié: (2023)
Stingy Context: 18:1 Hierarchical Code Compression for LLM Auto-Coding
par: Ostby, David Linus
Publié: (2026)
par: Ostby, David Linus
Publié: (2026)
Pragmatic Reasoning improves LLM Code Generation
par: Cao, Zhuchen, et autres
Publié: (2025)
par: Cao, Zhuchen, et autres
Publié: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
par: Tao, Tianhua, et autres
Publié: (2024)
par: Tao, Tianhua, et autres
Publié: (2024)
Let the Code LLM Edit Itself When You Edit the Code
par: He, Zhenyu, et autres
Publié: (2024)
par: He, Zhenyu, et autres
Publié: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
LocAgent: Graph-Guided LLM Agents for Code Localization
par: Chen, Zhaoling, et autres
Publié: (2025)
par: Chen, Zhaoling, et autres
Publié: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
par: Zhang, Ziyao, et autres
Publié: (2024)
par: Zhang, Ziyao, et autres
Publié: (2024)
LeDex: Training LLMs to Better Self-Debug and Explain Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
par: Ding, Yangruibo, et autres
Publié: (2024)
par: Ding, Yangruibo, et autres
Publié: (2024)
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
par: Arnaudo, Anna, et autres
Publié: (2026)
par: Arnaudo, Anna, et autres
Publié: (2026)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
par: Han, Hojae, et autres
Publié: (2025)
par: Han, Hojae, et autres
Publié: (2025)
Collaboration is all you need: LLM Assisted Safe Code Translation
par: Karanjai, Rabimba, et autres
Publié: (2025)
par: Karanjai, Rabimba, et autres
Publié: (2025)
Issue Localization via LLM-Driven Iterative Code Graph Searching
par: Jiang, Zhonghao, et autres
Publié: (2025)
par: Jiang, Zhonghao, et autres
Publié: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
par: Peng, Yun, et autres
Publié: (2024)
par: Peng, Yun, et autres
Publié: (2024)
Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code
par: Song, Zhenghan, et autres
Publié: (2026)
par: Song, Zhenghan, et autres
Publié: (2026)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
par: Paul, Indraneil, et autres
Publié: (2025)
par: Paul, Indraneil, et autres
Publié: (2025)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
par: Ficek, Aleksander, et autres
Publié: (2025)
par: Ficek, Aleksander, et autres
Publié: (2025)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
par: Gupta, Tanmay, et autres
Publié: (2024)
par: Gupta, Tanmay, et autres
Publié: (2024)
Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline
par: Zeng, Guancheng, et autres
Publié: (2024)
par: Zeng, Guancheng, et autres
Publié: (2024)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)
par: Cai, Songcheng, et autres
Publié: (2026)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
par: Chervyakov, Artem, et autres
Publié: (2025)
par: Chervyakov, Artem, et autres
Publié: (2025)
Dafny as Verification-Aware Intermediate Language for Code Generation
par: Li, Yue Chen, et autres
Publié: (2025)
par: Li, Yue Chen, et autres
Publié: (2025)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
par: Zhang, William, et autres
Publié: (2024)
par: Zhang, William, et autres
Publié: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
par: Tu, Xinming, et autres
Publié: (2026)
par: Tu, Xinming, et autres
Publié: (2026)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
par: Zhang, Tianyi, et autres
Publié: (2025)
par: Zhang, Tianyi, et autres
Publié: (2025)
Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries
par: Ullah, Muhammad Aziz, et autres
Publié: (2026)
par: Ullah, Muhammad Aziz, et autres
Publié: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
IndustryCode: A Benchmark for Industry Code Generation
par: Zeng, Puyu, et autres
Publié: (2026)
par: Zeng, Puyu, et autres
Publié: (2026)
Rethinking Code Refinement: Learning to Judge Code Efficiency
par: Seo, Minju, et autres
Publié: (2024)
par: Seo, Minju, et autres
Publié: (2024)
B4: Towards Optimal Assessment of Plausible Code Solutions with Plausible Tests
par: Chen, Mouxiang, et autres
Publié: (2024)
par: Chen, Mouxiang, et autres
Publié: (2024)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
par: Zhuo, Terry Yue, et autres
Publié: (2024)
par: Zhuo, Terry Yue, et autres
Publié: (2024)
A Code Comprehension Benchmark for Large Language Models for Code
par: Havare, Jayant, et autres
Publié: (2025)
par: Havare, Jayant, et autres
Publié: (2025)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
par: Zheng, Tianyu, et autres
Publié: (2024)
par: Zheng, Tianyu, et autres
Publié: (2024)
CodeMirage: Hallucinations in Code Generated by Large Language Models
par: Agarwal, Vibhor, et autres
Publié: (2024)
par: Agarwal, Vibhor, et autres
Publié: (2024)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
par: Zhao, Haoyu, et autres
Publié: (2026)
par: Zhao, Haoyu, et autres
Publié: (2026)
CodeS: Natural Language to Code Repository via Multi-Layer Sketch
par: Zan, Daoguang, et autres
Publié: (2024)
par: Zan, Daoguang, et autres
Publié: (2024)
Documents similaires
-
Pull Requests as a Training Signal for Repo-Level Code Editing
par: Zhu, Qinglin, et autres
Publié: (2026) -
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025) -
Training Versatile Coding Agents in Synthetic Environments
par: Zhu, Yiqi, et autres
Publié: (2025) -
Ranking LLM-Generated Loop Invariants for Program Verification
par: Chakraborty, Saikat, et autres
Publié: (2023) -
Stingy Context: 18:1 Hierarchical Code Compression for LLM Auto-Coding
par: Ostby, David Linus
Publié: (2026)