DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Enhao, Sun, Pengyu, Lin, Zixin, Chen, Alex, Ouyang, Joey, Wang, Haobo, Hu, Kaichun, Yi, James, Li, Frank, Zhang, Zhiyu, Xu, Tianxiang, Zhao, Gang, Ling, Ziang, Yang, Lowes |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DMind-3: A Sovereign Edge--Local--Cloud AI System with Controlled Deliberation and Correction-Based Tuning for Safe, Low-Latency Transaction Execution
by: Huang, Enhao, et al.
Published: (2026)
by: Huang, Enhao, et al.
Published: (2026)
Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
by: Huang, Enhao, et al.
Published: (2025)
by: Huang, Enhao, et al.
Published: (2025)
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
by: Yin, Guoli, et al.
Published: (2024)
by: Yin, Guoli, et al.
Published: (2024)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025)
by: Tian, Jiaming, et al.
Published: (2025)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Towards Cross-Table Masked Pretraining for Web Data Mining
by: Ye, Chao, et al.
Published: (2023)
by: Ye, Chao, et al.
Published: (2023)
DualMind: Towards Understanding Cognitive-Affective Cascades in Public Opinion Dissemination via Multi-Agent Simulation
by: Huang, Enhao, et al.
Published: (2026)
by: Huang, Enhao, et al.
Published: (2026)
LLM-Advisor: An LLM Benchmark for Cost-efficient Path Planning across Multiple Terrains
by: Xiao, Ling, et al.
Published: (2025)
by: Xiao, Ling, et al.
Published: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
by: Gu, Zhouhong, et al.
Published: (2023)
by: Gu, Zhouhong, et al.
Published: (2023)
Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
by: Bhardwaj, Arth, et al.
Published: (2026)
by: Bhardwaj, Arth, et al.
Published: (2026)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
by: Liu, Chuang, et al.
Published: (2024)
by: Liu, Chuang, et al.
Published: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
by: Pleines, Marco, et al.
Published: (2023)
by: Pleines, Marco, et al.
Published: (2023)
Moduli space of genus one curves on cubic threefold
by: Feng, Enhao
Published: (2025)
by: Feng, Enhao
Published: (2025)
MFF-EINV2: Multi-scale Feature Fusion across Spectral-Spatial-Temporal Domains for Sound Event Localization and Detection
by: Mu, Da, et al.
Published: (2024)
by: Mu, Da, et al.
Published: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)
by: Min, Hyangsuk, et al.
Published: (2025)
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities
by: Nishina, Kunato, et al.
Published: (2024)
by: Nishina, Kunato, et al.
Published: (2024)
LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
by: Zhou, Weilin, et al.
Published: (2026)
by: Zhou, Weilin, et al.
Published: (2026)
Across Time and (Product) Space: A Capability-Centric Model of Relatedness and Economic Complexity
by: Huang, Ziang, et al.
Published: (2025)
by: Huang, Ziang, et al.
Published: (2025)
Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model
by: Qi, Zelu, et al.
Published: (2025)
by: Qi, Zelu, et al.
Published: (2025)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
by: Hu, Tianxiang, et al.
Published: (2024)
by: Hu, Tianxiang, et al.
Published: (2024)
Holistic Energy Performance Management: Enablers, Capabilities, and Features
by: Masoudi, Meysam, et al.
Published: (2026)
by: Masoudi, Meysam, et al.
Published: (2026)
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
by: Jia, Hangyi, et al.
Published: (2025)
by: Jia, Hangyi, et al.
Published: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
by: Chen, Weiyi, et al.
Published: (2026)
by: Chen, Weiyi, et al.
Published: (2026)
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
by: Lin, Zhiyu, et al.
Published: (2025)
by: Lin, Zhiyu, et al.
Published: (2025)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
by: Liu, Junqi, et al.
Published: (2025)
by: Liu, Junqi, et al.
Published: (2025)
SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
by: Xiao, Zhiqing, et al.
Published: (2025)
by: Xiao, Zhiqing, et al.
Published: (2025)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
by: Zhang, Enci, et al.
Published: (2025)
by: Zhang, Enci, et al.
Published: (2025)
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
by: Zhu, Pengyu, et al.
Published: (2026)
by: Zhu, Pengyu, et al.
Published: (2026)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
by: Zhang, Yuanhan, et al.
Published: (2025)
by: Zhang, Yuanhan, et al.
Published: (2025)
Towards Holistic Prompt Craft
by: Lindley, Joseph, et al.
Published: (2025)
by: Lindley, Joseph, et al.
Published: (2025)
Similar Items
-
DMind-3: A Sovereign Edge--Local--Cloud AI System with Controlled Deliberation and Correction-Based Tuning for Safe, Low-Latency Transaction Execution
by: Huang, Enhao, et al.
Published: (2026) -
Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
by: Huang, Enhao, et al.
Published: (2025) -
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
by: Yin, Guoli, et al.
Published: (2024) -
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025) -
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
by: Ling Team, et al.
Published: (2025)