Learning Generalizable Multimodal Representations for Software Vulnerability Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dong, Zeming, Guo, Yuejun, Hu, Qiang, Zhang, Yao, Cordy, Maxime, Liu, Hao, Papadakis, Mike, Lyu, Yongqiang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911634815975424
author Dong, Zeming
Guo, Yuejun
Hu, Qiang
Zhang, Yao
Cordy, Maxime
Liu, Hao
Papadakis, Mike
Lyu, Yongqiang
author_facet Dong, Zeming
Guo, Yuejun
Hu, Qiang
Zhang, Yao
Cordy, Maxime
Liu, Hao
Papadakis, Mike
Lyu, Yongqiang
contents Source code and its accompanying comments are complementary yet naturally aligned modalities-code encodes structural logic while comments capture developer intent. However, existing vulnerability detection methods mostly rely on single-modality code representations, overlooking the complementary semantic information embedded in comments and thus limiting their generalization across complex code structures and logical relationships. To address this, we propose MultiVul, a multimodal contrastive framework that aligns code and comment representations through dual similarity learning and consistency regularization, augmented with diverse code-text pairs to improve robustness. Experiments on widely adopted DiverseVul and Devign datasets across four large language models (LLMs) (i.e., DeepSeek-Coder-6.7B, Qwen2.5-Coder-7B, StarCoder2-7B, and CodeLlama-7B) show that MultiVul achieves up to 27.07% F1 improvement over prompting-based methods and 13.37% over code-only Fine-Tuning, while maintaining comparable inference efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25711
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Generalizable Multimodal Representations for Software Vulnerability Detection
Dong, Zeming
Guo, Yuejun
Hu, Qiang
Zhang, Yao
Cordy, Maxime
Liu, Hao
Papadakis, Mike
Lyu, Yongqiang
Software Engineering
Artificial Intelligence
Source code and its accompanying comments are complementary yet naturally aligned modalities-code encodes structural logic while comments capture developer intent. However, existing vulnerability detection methods mostly rely on single-modality code representations, overlooking the complementary semantic information embedded in comments and thus limiting their generalization across complex code structures and logical relationships. To address this, we propose MultiVul, a multimodal contrastive framework that aligns code and comment representations through dual similarity learning and consistency regularization, augmented with diverse code-text pairs to improve robustness. Experiments on widely adopted DiverseVul and Devign datasets across four large language models (LLMs) (i.e., DeepSeek-Coder-6.7B, Qwen2.5-Coder-7B, StarCoder2-7B, and CodeLlama-7B) show that MultiVul achieves up to 27.07% F1 improvement over prompting-based methods and 13.37% over code-only Fine-Tuning, while maintaining comparable inference efficiency.
title Learning Generalizable Multimodal Representations for Software Vulnerability Detection
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2604.25711