Learning Generalizable Multimodal Representations for Software Vulnerability Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911634815975424 |
|---|---|
| author | Dong, Zeming Guo, Yuejun Hu, Qiang Zhang, Yao Cordy, Maxime Liu, Hao Papadakis, Mike Lyu, Yongqiang |
| author_facet | Dong, Zeming Guo, Yuejun Hu, Qiang Zhang, Yao Cordy, Maxime Liu, Hao Papadakis, Mike Lyu, Yongqiang |
| contents | Source code and its accompanying comments are complementary yet naturally aligned modalities-code encodes structural logic while comments capture developer intent. However, existing vulnerability detection methods mostly rely on single-modality code representations, overlooking the complementary semantic information embedded in comments and thus limiting their generalization across complex code structures and logical relationships. To address this, we propose MultiVul, a multimodal contrastive framework that aligns code and comment representations through dual similarity learning and consistency regularization, augmented with diverse code-text pairs to improve robustness. Experiments on widely adopted DiverseVul and Devign datasets across four large language models (LLMs) (i.e., DeepSeek-Coder-6.7B, Qwen2.5-Coder-7B, StarCoder2-7B, and CodeLlama-7B) show that MultiVul achieves up to 27.07% F1 improvement over prompting-based methods and 13.37% over code-only Fine-Tuning, while maintaining comparable inference efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_25711 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Learning Generalizable Multimodal Representations for Software Vulnerability Detection Dong, Zeming Guo, Yuejun Hu, Qiang Zhang, Yao Cordy, Maxime Liu, Hao Papadakis, Mike Lyu, Yongqiang Software Engineering Artificial Intelligence Source code and its accompanying comments are complementary yet naturally aligned modalities-code encodes structural logic while comments capture developer intent. However, existing vulnerability detection methods mostly rely on single-modality code representations, overlooking the complementary semantic information embedded in comments and thus limiting their generalization across complex code structures and logical relationships. To address this, we propose MultiVul, a multimodal contrastive framework that aligns code and comment representations through dual similarity learning and consistency regularization, augmented with diverse code-text pairs to improve robustness. Experiments on widely adopted DiverseVul and Devign datasets across four large language models (LLMs) (i.e., DeepSeek-Coder-6.7B, Qwen2.5-Coder-7B, StarCoder2-7B, and CodeLlama-7B) show that MultiVul achieves up to 27.07% F1 improvement over prompting-based methods and 13.37% over code-only Fine-Tuning, while maintaining comparable inference efficiency. |
| title | Learning Generalizable Multimodal Representations for Software Vulnerability Detection |
| topic | Software Engineering Artificial Intelligence |
| url | https://arxiv.org/abs/2604.25711 |