OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Lianghong, Tao, Wei, Jiang, Runhan, Wang, Yanlin, Chen, Jiachi, Liu, Xilin, Ma, Yuchi, Mao, Mingzhi, Zhang, Hongyu, Zheng, Zibin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913825711718400
author Guo, Lianghong
Tao, Wei
Jiang, Runhan
Wang, Yanlin
Chen, Jiachi
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
author_facet Guo, Lianghong
Tao, Wei
Jiang, Runhan
Wang, Yanlin
Chen, Jiachi
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
contents The GitHub issue resolution task aims to resolve issues reported in repositories automatically. With advances in large language models (LLMs), this task has gained increasing attention, and several benchmarks are proposed to evaluate the issue resolution ability of LLMs. However, existing benchmarks have three main limitations. First, current benchmarks focus on a single programming language, limiting the evaluation of issues from repositories across different languages. Second, they usually cover a narrow range of domains, which may fail to represent the diversity of real-world issues. Third, existing benchmarks rely solely on textual information in issue descriptions, overlooking multimodal information such as images in issues. In this paper, we propose OmniGIRL, a GitHub Issue ResoLution benchmark that is multilingual, multimodal, and multi-domain. OmniGIRL includes 959 task instances, which are collected from repositories across four programming languages (i.e., Python, JavaScript, TypeScript, and Java) and eight different domains. Our evaluation shows that current LLMs show limited performances on OmniGIRL. Notably, the best-performing model, GPT-4o, resolves only 8.6% of the issues. Besides, we find that current LLMs struggle to resolve issues requiring understanding images. The best performance is achieved by Claude-3.5-Sonnet, which resolves only 10.5% of the issues with image information. Finally, we analyze the reasons behind current LLMs' failure on OmniGIRL, providing insights for future improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
Guo, Lianghong
Tao, Wei
Jiang, Runhan
Wang, Yanlin
Chen, Jiachi
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
Software Engineering
The GitHub issue resolution task aims to resolve issues reported in repositories automatically. With advances in large language models (LLMs), this task has gained increasing attention, and several benchmarks are proposed to evaluate the issue resolution ability of LLMs. However, existing benchmarks have three main limitations. First, current benchmarks focus on a single programming language, limiting the evaluation of issues from repositories across different languages. Second, they usually cover a narrow range of domains, which may fail to represent the diversity of real-world issues. Third, existing benchmarks rely solely on textual information in issue descriptions, overlooking multimodal information such as images in issues. In this paper, we propose OmniGIRL, a GitHub Issue ResoLution benchmark that is multilingual, multimodal, and multi-domain. OmniGIRL includes 959 task instances, which are collected from repositories across four programming languages (i.e., Python, JavaScript, TypeScript, and Java) and eight different domains. Our evaluation shows that current LLMs show limited performances on OmniGIRL. Notably, the best-performing model, GPT-4o, resolves only 8.6% of the issues. Besides, we find that current LLMs struggle to resolve issues requiring understanding images. The best performance is achieved by Claude-3.5-Sonnet, which resolves only 10.5% of the issues with image information. Finally, we analyze the reasons behind current LLMs' failure on OmniGIRL, providing insights for future improvements.
title OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
topic Software Engineering
url https://arxiv.org/abs/2505.04606