An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gandhi, Shubham, Naik, Atharva, Xie, Yiqing, Rose, Carolyn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
by: Naik, Atharva
Published: (2024)
by: Naik, Atharva
Published: (2024)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024)
by: Naik, Atharva, et al.
Published: (2024)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
by: Xie, Yiqing, et al.
Published: (2023)
by: Xie, Yiqing, et al.
Published: (2023)
Neurosymbolic Repo-level Code Localization
by: Xu, Xiufeng, et al.
Published: (2026)
by: Xu, Xiufeng, et al.
Published: (2026)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
by: Xie, Yiqing, et al.
Published: (2025)
by: Xie, Yiqing, et al.
Published: (2025)
Contract-Coding: Towards Repo-Level Generation via Structured Symbolic Paradigm
by: Lin, Yi, et al.
Published: (2026)
by: Lin, Yi, et al.
Published: (2026)
MetaLint: Easy-to-Hard Generalization for Code Linting
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
by: Wu, Qinyun, et al.
Published: (2024)
by: Wu, Qinyun, et al.
Published: (2024)
RepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
by: Pan, Zhongqiang, et al.
Published: (2026)
by: Pan, Zhongqiang, et al.
Published: (2026)
Tool-integrated Reinforcement Learning for Repo Deep Search
by: Ma, Zexiong, et al.
Published: (2025)
by: Ma, Zexiong, et al.
Published: (2025)
RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
by: Ouyang, Siru, et al.
Published: (2024)
by: Ouyang, Siru, et al.
Published: (2024)
RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion
by: Phan, Huy N., et al.
Published: (2024)
by: Phan, Huy N., et al.
Published: (2024)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
by: Hu, Ruida, et al.
Published: (2024)
by: Hu, Ruida, et al.
Published: (2024)
Beyond Autoregression: An Empirical Study of Diffusion Large Language Models for Code Generation
by: Li, Chengze, et al.
Published: (2025)
by: Li, Chengze, et al.
Published: (2025)
RepoReviewer: A Local-First Multi-Agent Architecture for Repository-Level Code Review
by: Zhang, Peng
Published: (2026)
by: Zhang, Peng
Published: (2026)
Needle in the Repo: A Benchmark for Maintainability in AI-Generated Repository Edits
by: Zhu, Haichao, et al.
Published: (2026)
by: Zhu, Haichao, et al.
Published: (2026)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
by: Vartziotis, Tina, et al.
Published: (2024)
by: Vartziotis, Tina, et al.
Published: (2024)
Comparison of Static Application Security Testing Tools and Large Language Models for Repo-level Vulnerability Detection
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
An Empirical Study on Self-correcting Large Language Models for Data Science Code Generation
by: Quoc, Thai Tang, et al.
Published: (2024)
by: Quoc, Thai Tang, et al.
Published: (2024)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
Pull Requests as a Training Signal for Repo-Level Code Editing
by: Zhu, Qinglin, et al.
Published: (2026)
by: Zhu, Qinglin, et al.
Published: (2026)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
by: Nguyen, Thu-Trang, et al.
Published: (2024)
by: Nguyen, Thu-Trang, et al.
Published: (2024)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
How Do Agents Perform Code Optimization? An Empirical Study
by: Peng, Huiyun, et al.
Published: (2025)
by: Peng, Huiyun, et al.
Published: (2025)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
Empirical Studies of Parameter Efficient Methods for Large Language Models of Code and Knowledge Transfer to R
by: Esmaeili, Amirreza, et al.
Published: (2024)
by: Esmaeili, Amirreza, et al.
Published: (2024)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
by: Majgaonkar, Oorja, et al.
Published: (2025)
by: Majgaonkar, Oorja, et al.
Published: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
by: Wang, Huacan, et al.
Published: (2025)
by: Wang, Huacan, et al.
Published: (2025)
Designing Empirical Studies on LLM-Based Code Generation: Towards a Reference Framework
by: Nascimento, Nathalia, et al.
Published: (2025)
by: Nascimento, Nathalia, et al.
Published: (2025)
Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study
by: Bansal, Kaushal
Published: (2026)
by: Bansal, Kaushal
Published: (2026)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
by: Dong, Zeming, et al.
Published: (2023)
by: Dong, Zeming, et al.
Published: (2023)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
by: Geruslu, Vehid, et al.
Published: (2026)
by: Geruslu, Vehid, et al.
Published: (2026)
Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
by: Della Porta, Antonio, et al.
Published: (2025)
by: Della Porta, Antonio, et al.
Published: (2025)
Similar Items
-
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
by: Naik, Atharva
Published: (2024) -
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024) -
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
by: Xie, Yiqing, et al.
Published: (2023) -
Neurosymbolic Repo-level Code Localization
by: Xu, Xiufeng, et al.
Published: (2026) -
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)