AICrypto: Evaluating Cryptography Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yu, Liu, Yijian, Ji, Liheng, Luo, Han, Li, Wenjie, Zhou, Xiaofei, Feng, Chiyun, Wang, Puji, Cao, Yuhan, Zhang, Geyuan, Li, Xiaojian, Xu, Rongwu, Chen, Yilei, He, Tianxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913167498543104
author Wang, Yu
Liu, Yijian
Ji, Liheng
Luo, Han
Li, Wenjie
Zhou, Xiaofei
Feng, Chiyun
Wang, Puji
Cao, Yuhan
Zhang, Geyuan
Li, Xiaojian
Xu, Rongwu
Chen, Yilei
He, Tianxing
author_facet Wang, Yu
Liu, Yijian
Ji, Liheng
Luo, Han
Li, Wenjie
Zhou, Xiaofei
Feng, Chiyun
Wang, Puji
Cao, Yuhan
Zhang, Geyuan
Li, Xiaojian
Xu, Rongwu
Chen, Yilei
He, Tianxing
contents We build \textbf{AICrypto}, a comprehensive benchmark designed to evaluate the cryptography capabilities of large language models (LLMs). The benchmark comprises 135 multiple-choice questions, 150 capture-the-flag challenges, and 30 proof problems, covering a broad range of skills from knowledge memorization to vulnerability exploitation and formal reasoning. All tasks are carefully reviewed or constructed by cryptography experts to improve correctness and rigor. For each proof problem, we provide detailed scoring rubrics and reference solutions that enable automated grading, achieving high correlation with human expert evaluations. We introduce strong human expert performance baselines for comparison across all task types. Our evaluation of 17 leading LLMs reveals that state-of-the-art models match or even surpass human experts in memorizing cryptographic concepts, exploiting common vulnerabilities, and routine proofs. However, our analysis reveals that they still lack a deep understanding of abstract mathematical concepts and struggle with tasks that require multi-step reasoning and dynamic analysis. We hope this work could provide insights for future research on LLMs in cryptographic applications. Our code and dataset are available at https://github.com/wangyu-ovo/aicrypto-agent.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09580
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AICrypto: Evaluating Cryptography Capabilities of Large Language Models
Wang, Yu
Liu, Yijian
Ji, Liheng
Luo, Han
Li, Wenjie
Zhou, Xiaofei
Feng, Chiyun
Wang, Puji
Cao, Yuhan
Zhang, Geyuan
Li, Xiaojian
Xu, Rongwu
Chen, Yilei
He, Tianxing
Cryptography and Security
We build \textbf{AICrypto}, a comprehensive benchmark designed to evaluate the cryptography capabilities of large language models (LLMs). The benchmark comprises 135 multiple-choice questions, 150 capture-the-flag challenges, and 30 proof problems, covering a broad range of skills from knowledge memorization to vulnerability exploitation and formal reasoning. All tasks are carefully reviewed or constructed by cryptography experts to improve correctness and rigor. For each proof problem, we provide detailed scoring rubrics and reference solutions that enable automated grading, achieving high correlation with human expert evaluations. We introduce strong human expert performance baselines for comparison across all task types. Our evaluation of 17 leading LLMs reveals that state-of-the-art models match or even surpass human experts in memorizing cryptographic concepts, exploiting common vulnerabilities, and routine proofs. However, our analysis reveals that they still lack a deep understanding of abstract mathematical concepts and struggle with tasks that require multi-step reasoning and dynamic analysis. We hope this work could provide insights for future research on LLMs in cryptographic applications. Our code and dataset are available at https://github.com/wangyu-ovo/aicrypto-agent.
title AICrypto: Evaluating Cryptography Capabilities of Large Language Models
topic Cryptography and Security
url https://arxiv.org/abs/2507.09580