OCEAN: Open-World Contrastive Authorship Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mächtle, Felix, Serr, Jan-Niclas, Loose, Nils, Sander, Jonas, Eisenbarth, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913600106397696
author Mächtle, Felix
Serr, Jan-Niclas
Loose, Nils
Sander, Jonas
Eisenbarth, Thomas
author_facet Mächtle, Felix
Serr, Jan-Niclas
Loose, Nils
Sander, Jonas
Eisenbarth, Thomas
contents In an era where cyberattacks increasingly target the software supply chain, the ability to accurately attribute code authorship in binary files is critical to improving cybersecurity measures. We propose OCEAN, a contrastive learning-based system for function-level authorship attribution. OCEAN is the first framework to explore code authorship attribution on compiled binaries in an open-world and extreme scenario, where two code samples from unknown authors are compared to determine if they are developed by the same author. To evaluate OCEAN, we introduce new realistic datasets: CONAN, to improve the performance of authorship attribution systems in real-world use cases, and SNOOPY, to increase the robustness of the evaluation of such systems. We use CONAN to train our model and evaluate on SNOOPY, a fully unseen dataset, resulting in an AUROC score of 0.86 even when using high compiler optimizations. We further show that CONAN improves performance by 7% compared to the previously used Google Code Jam dataset. Additionally, OCEAN outperforms previous methods in their settings, achieving a 10% improvement over state-of-the-art SCS-Gan in scenarios analyzing source code. Furthermore, OCEAN can detect code injections from an unknown author in a software update, underscoring its value for securing software supply chains.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OCEAN: Open-World Contrastive Authorship Identification
Mächtle, Felix
Serr, Jan-Niclas
Loose, Nils
Sander, Jonas
Eisenbarth, Thomas
Artificial Intelligence
Cryptography and Security
In an era where cyberattacks increasingly target the software supply chain, the ability to accurately attribute code authorship in binary files is critical to improving cybersecurity measures. We propose OCEAN, a contrastive learning-based system for function-level authorship attribution. OCEAN is the first framework to explore code authorship attribution on compiled binaries in an open-world and extreme scenario, where two code samples from unknown authors are compared to determine if they are developed by the same author. To evaluate OCEAN, we introduce new realistic datasets: CONAN, to improve the performance of authorship attribution systems in real-world use cases, and SNOOPY, to increase the robustness of the evaluation of such systems. We use CONAN to train our model and evaluate on SNOOPY, a fully unseen dataset, resulting in an AUROC score of 0.86 even when using high compiler optimizations. We further show that CONAN improves performance by 7% compared to the previously used Google Code Jam dataset. Additionally, OCEAN outperforms previous methods in their settings, achieving a 10% improvement over state-of-the-art SCS-Gan in scenarios analyzing source code. Furthermore, OCEAN can detect code injections from an unknown author in a software update, underscoring its value for securing software supply chains.
title OCEAN: Open-World Contrastive Authorship Identification
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2412.05049