CodeV: Issue Resolving with Visual Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Linhao, Zan, Daoguang, Yang, Quanshun, Huang, Zhirong, Chen, Dong, Shen, Bo, Liu, Tianyu, Gong, Yongshun, Huang, Pengjie, Lu, Xudong, Liang, Guangtai, Cui, Lizhen, Wang, Qianxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910759645085696
author Zhang, Linhao
Zan, Daoguang
Yang, Quanshun
Huang, Zhirong
Chen, Dong
Shen, Bo
Liu, Tianyu
Gong, Yongshun
Huang, Pengjie
Lu, Xudong
Liang, Guangtai
Cui, Lizhen
Wang, Qianxiang
author_facet Zhang, Linhao
Zan, Daoguang
Yang, Quanshun
Huang, Zhirong
Chen, Dong
Shen, Bo
Liu, Tianyu
Gong, Yongshun
Huang, Pengjie
Lu, Xudong
Liang, Guangtai
Cui, Lizhen
Wang, Qianxiang
contents Large Language Models (LLMs) have advanced rapidly in recent years, with their applications in software engineering expanding to more complex repository-level tasks. GitHub issue resolving is a key challenge among these tasks. While recent approaches have made progress on this task, they focus on textual data within issues, neglecting visual data. However, this visual data is crucial for resolving issues as it conveys additional knowledge that text alone cannot. We propose CodeV, the first approach to leveraging visual data to enhance the issue-resolving capabilities of LLMs. CodeV resolves each issue by following a two-phase process: data processing and patch generation. To evaluate CodeV, we construct a benchmark for visual issue resolving, namely Visual SWE-bench. Through extensive experiments, we demonstrate the effectiveness of CodeV, as well as provide valuable insights into leveraging visual data to resolve GitHub issues.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeV: Issue Resolving with Visual Data
Zhang, Linhao
Zan, Daoguang
Yang, Quanshun
Huang, Zhirong
Chen, Dong
Shen, Bo
Liu, Tianyu
Gong, Yongshun
Huang, Pengjie
Lu, Xudong
Liang, Guangtai
Cui, Lizhen
Wang, Qianxiang
Software Engineering
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have advanced rapidly in recent years, with their applications in software engineering expanding to more complex repository-level tasks. GitHub issue resolving is a key challenge among these tasks. While recent approaches have made progress on this task, they focus on textual data within issues, neglecting visual data. However, this visual data is crucial for resolving issues as it conveys additional knowledge that text alone cannot. We propose CodeV, the first approach to leveraging visual data to enhance the issue-resolving capabilities of LLMs. CodeV resolves each issue by following a two-phase process: data processing and patch generation. To evaluate CodeV, we construct a benchmark for visual issue resolving, namely Visual SWE-bench. Through extensive experiments, we demonstrate the effectiveness of CodeV, as well as provide valuable insights into leveraging visual data to resolve GitHub issues.
title CodeV: Issue Resolving with Visual Data
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.17315