Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liyanage, Danushka, Attanayake, Nelum, Luo, Zijian, Gopinath, Rahul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908462111260672
author Liyanage, Danushka
Attanayake, Nelum
Luo, Zijian
Gopinath, Rahul
author_facet Liyanage, Danushka
Attanayake, Nelum
Luo, Zijian
Gopinath, Rahul
contents Background: Fuzzers are often guided by coverage, making the estimation of maximum achievable coverage a key concern in fuzzing. However, achieving 100% coverage is infeasible for most real-world software systems, regardless of effort. While static reachability analysis can provide an upper bound, it is often highly inaccurate. Recently, statistical estimation methods based on species richness estimators from biostatistics have been proposed as a potential solution. Yet, the lack of reliable benchmarks with labeled ground truth has limited rigorous evaluation of their accuracy. Objective: This work examines the reliability of reachability estimators from two axes: addressing the lack of labeled ground truth and evaluating their reliability on real-world programs. Methods: (1) To address the challenge of labeled ground truth, we propose an evaluation framework that synthetically generates large programs with complex control flows, ensuring well-defined reachability and providing ground truth for evaluation. (2) To address the criticism from use of synthetic benchmarks, we adapt a reliability check for reachability estimators on real-world benchmarks without labeled ground truth -- by varying the size of sampling units, which, in theory, should not affect the estimate. Results: These two studies together will help answer the question of whether current reachability estimators are reliable, and defines a protocol to evaluate future improvements in reachability estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
Liyanage, Danushka
Attanayake, Nelum
Luo, Zijian
Gopinath, Rahul
Software Engineering
68N30
D.2.4; D.2.5; D.2.8
Background: Fuzzers are often guided by coverage, making the estimation of maximum achievable coverage a key concern in fuzzing. However, achieving 100% coverage is infeasible for most real-world software systems, regardless of effort. While static reachability analysis can provide an upper bound, it is often highly inaccurate. Recently, statistical estimation methods based on species richness estimators from biostatistics have been proposed as a potential solution. Yet, the lack of reliable benchmarks with labeled ground truth has limited rigorous evaluation of their accuracy. Objective: This work examines the reliability of reachability estimators from two axes: addressing the lack of labeled ground truth and evaluating their reliability on real-world programs. Methods: (1) To address the challenge of labeled ground truth, we propose an evaluation framework that synthetically generates large programs with complex control flows, ensuring well-defined reachability and providing ground truth for evaluation. (2) To address the criticism from use of synthetic benchmarks, we adapt a reliability check for reachability estimators on real-world benchmarks without labeled ground truth -- by varying the size of sampling units, which, in theory, should not affect the estimate. Results: These two studies together will help answer the question of whether current reachability estimators are reliable, and defines a protocol to evaluate future improvements in reachability estimation.
title Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
topic Software Engineering
68N30
D.2.4; D.2.5; D.2.8
url https://arxiv.org/abs/2507.17093