KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Xue, Li, Ge, Qian, Jiaru, Shi, Xianjie, Li, Chenjie, Zhu, Hao, Wang, Ziyu, Zhang, Jielun, Zhao, Zheyu, Wu, Lingwei, Zhang, Kechi, Li, Jia, Jiao, Wenpin, Jin, Zhi, Dong, Yihong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917435871854592
author Jiang, Xue
Li, Ge
Qian, Jiaru
Shi, Xianjie
Li, Chenjie
Zhu, Hao
Wang, Ziyu
Zhang, Jielun
Zhao, Zheyu
Wu, Lingwei
Zhang, Kechi
Li, Jia
Jiao, Wenpin
Jin, Zhi
Dong, Yihong
author_facet Jiang, Xue
Li, Ge
Qian, Jiaru
Shi, Xianjie
Li, Chenjie
Zhu, Hao
Wang, Ziyu
Zhang, Jielun
Zhao, Zheyu
Wu, Lingwei
Zhang, Kechi
Li, Jia
Jiao, Wenpin
Jin, Zhi
Dong, Yihong
contents Large language models (LLMs) excel at general programming but struggle with domain-specific software development, necessitating domain specialization methods for LLMs to learn and utilize domain knowledge and data. However, existing domain-specific code benchmarks cannot evaluate the effectiveness of domain specialization methods, which focus on assessing what knowledge LLMs possess rather than how they acquire and apply new knowledge, lacking explicit knowledge corpora for developing domain specialization methods. To this end, we present KOCO-BENCH, a novel benchmark designed for evaluating domain specialization methods in real-world software development. KOCO-BENCH contains 6 emerging domains with 11 software frameworks and 25 projects, featuring curated knowledge corpora alongside multi-granularity evaluation tasks including domain code generation (from function-level to project-level with rigorous test suites) and domain knowledge understanding (via multiple-choice Q&A). Unlike previous benchmarks that only provide test sets for direct evaluation, KOCO-BENCH requires acquiring and applying diverse domain knowledge (APIs, rules, constraints, etc.) from knowledge corpora to solve evaluation tasks. Our evaluations reveal that KOCO-BENCH poses significant challenges to state-of-the-art LLMs. Even with domain specialization methods (e.g., SFT, RAG, kNN-LM) applied, improvements remain marginal. Best-performing coding agent, Claude Code, achieves only 34.2%, highlighting the urgent need for more effective domain specialization methods. We release KOCO-BENCH, evaluation code, and baselines to advance further research at https://github.com/jiangxxxue/KOCO-bench.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13240
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?
Jiang, Xue
Li, Ge
Qian, Jiaru
Shi, Xianjie
Li, Chenjie
Zhu, Hao
Wang, Ziyu
Zhang, Jielun
Zhao, Zheyu
Wu, Lingwei
Zhang, Kechi
Li, Jia
Jiao, Wenpin
Jin, Zhi
Dong, Yihong
Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
Large language models (LLMs) excel at general programming but struggle with domain-specific software development, necessitating domain specialization methods for LLMs to learn and utilize domain knowledge and data. However, existing domain-specific code benchmarks cannot evaluate the effectiveness of domain specialization methods, which focus on assessing what knowledge LLMs possess rather than how they acquire and apply new knowledge, lacking explicit knowledge corpora for developing domain specialization methods. To this end, we present KOCO-BENCH, a novel benchmark designed for evaluating domain specialization methods in real-world software development. KOCO-BENCH contains 6 emerging domains with 11 software frameworks and 25 projects, featuring curated knowledge corpora alongside multi-granularity evaluation tasks including domain code generation (from function-level to project-level with rigorous test suites) and domain knowledge understanding (via multiple-choice Q&A). Unlike previous benchmarks that only provide test sets for direct evaluation, KOCO-BENCH requires acquiring and applying diverse domain knowledge (APIs, rules, constraints, etc.) from knowledge corpora to solve evaluation tasks. Our evaluations reveal that KOCO-BENCH poses significant challenges to state-of-the-art LLMs. Even with domain specialization methods (e.g., SFT, RAG, kNN-LM) applied, improvements remain marginal. Best-performing coding agent, Claude Code, achieves only 34.2%, highlighting the urgent need for more effective domain specialization methods. We release KOCO-BENCH, evaluation code, and baselines to advance further research at https://github.com/jiangxxxue/KOCO-bench.
title KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?
topic Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.13240