C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sirlanci, Melih, Yagemann, Carter, Lin, Zhiqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916699607924736
author Sirlanci, Melih
Yagemann, Carter
Lin, Zhiqiang
author_facet Sirlanci, Melih
Yagemann, Carter
Lin, Zhiqiang
contents Despite the effort in vulnerability detection over the last two decades, memory safety vulnerabilities continue to be a critical problem. Recent reports suggest that the key solution is to migrate to memory-safe languages. To this end, C-to-Rust transpilation becomes popular to resolve memory-safety issues in C programs. Recent works propose C-to-Rust transpilation frameworks; however, a comprehensive evaluation dataset is missing. Although one solution is to put together a large enough dataset, this increases the analysis time in automated frameworks as well as in manual efforts for some cases. In this work, we build a method to select functions from a large set to construct a minimized yet representative dataset to evaluate the C-to-Rust transpilation. We propose C2RUST-BENCH that contains 2,905 functions, which are representative of C-to-Rust transpilation, selected from 15,503 functions of real-world programs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation
Sirlanci, Melih
Yagemann, Carter
Lin, Zhiqiang
Cryptography and Security
Artificial Intelligence
Programming Languages
Despite the effort in vulnerability detection over the last two decades, memory safety vulnerabilities continue to be a critical problem. Recent reports suggest that the key solution is to migrate to memory-safe languages. To this end, C-to-Rust transpilation becomes popular to resolve memory-safety issues in C programs. Recent works propose C-to-Rust transpilation frameworks; however, a comprehensive evaluation dataset is missing. Although one solution is to put together a large enough dataset, this increases the analysis time in automated frameworks as well as in manual efforts for some cases. In this work, we build a method to select functions from a large set to construct a minimized yet representative dataset to evaluate the C-to-Rust transpilation. We propose C2RUST-BENCH that contains 2,905 functions, which are representative of C-to-Rust transpilation, selected from 15,503 functions of real-world programs.
title C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation
topic Cryptography and Security
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2504.15144