DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Enhao, Sun, Pengyu, Lin, Zixin, Chen, Alex, Ouyang, Joey, Wang, Haobo, Hu, Kaichun, Yi, James, Li, Frank, Zhang, Zhiyu, Xu, Tianxiang, Zhao, Gang, Ling, Ziang, Yang, Lowes
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911248374824960
author Huang, Enhao
Sun, Pengyu
Lin, Zixin
Chen, Alex
Ouyang, Joey
Wang, Haobo
Hu, Kaichun
Yi, James
Li, Frank
Zhang, Zhiyu
Xu, Tianxiang
Zhao, Gang
Ling, Ziang
Yang, Lowes
author_facet Huang, Enhao
Sun, Pengyu
Lin, Zixin
Chen, Alex
Ouyang, Joey
Wang, Haobo
Hu, Kaichun
Yi, James
Li, Frank
Zhang, Zhiyu
Xu, Tianxiang
Zhao, Gang
Ling, Ziang
Yang, Lowes
contents Large Language Models (LLMs) have achieved impressive performance in diverse natural language processing tasks, but specialized domains such as Web3 present new challenges and require more tailored evaluation. Despite the significant user base and capital flows in Web3, encompassing smart contracts, decentralized finance (DeFi), non-fungible tokens (NFTs), decentralized autonomous organizations (DAOs), on-chain governance, and novel token-economics, no comprehensive benchmark has systematically assessed LLM performance in this domain. To address this gap, we introduce the DMind Benchmark, a holistic Web3-oriented evaluation suite covering nine critical subfields: fundamental blockchain concepts, blockchain infrastructure, smart contract, DeFi mechanisms, DAOs, NFTs, token economics, meme concept, and security vulnerabilities. Beyond multiple-choice questions, DMind Benchmark features domain-specific tasks such as contract debugging and on-chain numeric reasoning, mirroring real-world scenarios. We evaluated 26 models, including ChatGPT, Claude, DeepSeek, Gemini, Grok, and Qwen, uncovering notable performance gaps in specialized areas like token economics and security-critical contract analysis. While some models excel in blockchain infrastructure tasks, advanced subfields remain challenging. Our benchmark dataset and evaluation pipeline are open-sourced on https://huggingface.co/datasets/DMindAI/DMind_Benchmark, reaching number one in Hugging Face's trending dataset charts within a week of release.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
Huang, Enhao
Sun, Pengyu
Lin, Zixin
Chen, Alex
Ouyang, Joey
Wang, Haobo
Hu, Kaichun
Yi, James
Li, Frank
Zhang, Zhiyu
Xu, Tianxiang
Zhao, Gang
Ling, Ziang
Yang, Lowes
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) have achieved impressive performance in diverse natural language processing tasks, but specialized domains such as Web3 present new challenges and require more tailored evaluation. Despite the significant user base and capital flows in Web3, encompassing smart contracts, decentralized finance (DeFi), non-fungible tokens (NFTs), decentralized autonomous organizations (DAOs), on-chain governance, and novel token-economics, no comprehensive benchmark has systematically assessed LLM performance in this domain. To address this gap, we introduce the DMind Benchmark, a holistic Web3-oriented evaluation suite covering nine critical subfields: fundamental blockchain concepts, blockchain infrastructure, smart contract, DeFi mechanisms, DAOs, NFTs, token economics, meme concept, and security vulnerabilities. Beyond multiple-choice questions, DMind Benchmark features domain-specific tasks such as contract debugging and on-chain numeric reasoning, mirroring real-world scenarios. We evaluated 26 models, including ChatGPT, Claude, DeepSeek, Gemini, Grok, and Qwen, uncovering notable performance gaps in specialized areas like token economics and security-critical contract analysis. While some models excel in blockchain infrastructure tasks, advanced subfields remain challenging. Our benchmark dataset and evaluation pipeline are open-sourced on https://huggingface.co/datasets/DMindAI/DMind_Benchmark, reaching number one in Hugging Face's trending dataset charts within a week of release.
title DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2504.16116