EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fan, Yiming, Won, Jun Yeon, Zhu, Ding, Sirlanci, Melih, Khalili, Mahdi, Yagemann, Carter
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908933126356992
author Fan, Yiming
Won, Jun Yeon
Zhu, Ding
Sirlanci, Melih
Khalili, Mahdi
Yagemann, Carter
author_facet Fan, Yiming
Won, Jun Yeon
Zhu, Ding
Sirlanci, Melih
Khalili, Mahdi
Yagemann, Carter
contents Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this application; however, due to the lack of a comprehensive universal benchmark in this field, researchers have struggled to compare different models effectively. Existing datasets are limited in scope, often focusing on a narrow set of transformations or types of binaries, and fail to reflect the full diversity of real-world applications. We introduce EXHIB, a benchmark comprising five realistic datasets collected from the wild, each highlighting a distinct aspect of the BFSD problem space. We evaluate 9 representative models spanning multiple BFSD paradigms on EXHIB and observe performance degradations of up to 30% on firmware and semantic datasets compared to standard settings, revealing substantial generalization gaps. Our results show that robustness to low- and mid-level binary variations does not generalize to high-level semantic differences, underscoring a critical blind spot in current BFSD evaluation practices.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01554
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild
Fan, Yiming
Won, Jun Yeon
Zhu, Ding
Sirlanci, Melih
Khalili, Mahdi
Yagemann, Carter
Cryptography and Security
Machine Learning
Software Engineering
D.4.6; K.6.5
Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this application; however, due to the lack of a comprehensive universal benchmark in this field, researchers have struggled to compare different models effectively. Existing datasets are limited in scope, often focusing on a narrow set of transformations or types of binaries, and fail to reflect the full diversity of real-world applications. We introduce EXHIB, a benchmark comprising five realistic datasets collected from the wild, each highlighting a distinct aspect of the BFSD problem space. We evaluate 9 representative models spanning multiple BFSD paradigms on EXHIB and observe performance degradations of up to 30% on firmware and semantic datasets compared to standard settings, revealing substantial generalization gaps. Our results show that robustness to low- and mid-level binary variations does not generalize to high-level semantic differences, underscoring a critical blind spot in current BFSD evaluation practices.
title EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild
topic Cryptography and Security
Machine Learning
Software Engineering
D.4.6; K.6.5
url https://arxiv.org/abs/2604.01554