CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Zhihao, Ji, Zimo, Zheng, Tao, Ren, Hao, Lan, Xiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909738103472128
author Li, Zhihao
Ji, Zimo
Zheng, Tao
Ren, Hao
Lan, Xiao
author_facet Li, Zhihao
Ji, Zimo
Zheng, Tao
Ren, Hao
Lan, Xiao
contents Cryptographic algorithms are fundamental to modern security, yet their implementations frequently harbor subtle logic flaws that are hard to detect. We introduce CryptoScope, a novel framework for automated cryptographic vulnerability detection powered by Large Language Models (LLMs). CryptoScope combines Chain-of-Thought (CoT) prompting with Retrieval-Augmented Generation (RAG), guided by a curated cryptographic knowledge base containing over 12,000 entries. We evaluate CryptoScope on LLM-CLVA, a benchmark of 92 cases primarily derived from real-world CVE vulnerabilities, complemented by cryptographic challenges from major Capture The Flag (CTF) competitions and synthetic examples across 11 programming languages. CryptoScope consistently improves performance over strong LLM baselines, boosting DeepSeek-V3 by 11.62%, GPT-4o-mini by 20.28%, and GLM-4-Flash by 28.69%. Additionally, it identifies 9 previously undisclosed flaws in widely used open-source cryptographic projects.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
Li, Zhihao
Ji, Zimo
Zheng, Tao
Ren, Hao
Lan, Xiao
Cryptography and Security
Artificial Intelligence
Cryptographic algorithms are fundamental to modern security, yet their implementations frequently harbor subtle logic flaws that are hard to detect. We introduce CryptoScope, a novel framework for automated cryptographic vulnerability detection powered by Large Language Models (LLMs). CryptoScope combines Chain-of-Thought (CoT) prompting with Retrieval-Augmented Generation (RAG), guided by a curated cryptographic knowledge base containing over 12,000 entries. We evaluate CryptoScope on LLM-CLVA, a benchmark of 92 cases primarily derived from real-world CVE vulnerabilities, complemented by cryptographic challenges from major Capture The Flag (CTF) competitions and synthetic examples across 11 programming languages. CryptoScope consistently improves performance over strong LLM baselines, boosting DeepSeek-V3 by 11.62%, GPT-4o-mini by 20.28%, and GLM-4-Flash by 28.69%. Additionally, it identifies 9 previously undisclosed flaws in widely used open-source cryptographic projects.
title CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2508.11599