Saved in:
| Main Authors: | Zhu, Longyuan, Hua, Hairan, Miao, Linlin, Zhao, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.11674 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026)
by: Sui, Xingyu, et al.
Published: (2026)
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
by: Akhtar, Mubashara, et al.
Published: (2026)
by: Akhtar, Mubashara, et al.
Published: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
by: Xu, Chengliang, et al.
Published: (2026)
by: Xu, Chengliang, et al.
Published: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
Benchmarking Overton Pluralism in LLMs
by: Poole-Dayan, Elinor, et al.
Published: (2025)
by: Poole-Dayan, Elinor, et al.
Published: (2025)
PHM-Bench: A Domain-Specific Benchmarking Framework for Systematic Evaluation of Large Models in Prognostics and Health Management
by: Yang, Puyu, et al.
Published: (2025)
by: Yang, Puyu, et al.
Published: (2025)
VCBench: Benchmarking LLMs in Venture Capital
by: Chen, Rick, et al.
Published: (2025)
by: Chen, Rick, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
LLMTM: Benchmarking and Optimizing LLMs for Temporal Motif Analysis in Dynamic Graphs
by: Hao, Bing, et al.
Published: (2025)
by: Hao, Bing, et al.
Published: (2025)
Fluidity Index: Next-Generation Super-intelligence Benchmarks
by: Ngoiya, Eric, et al.
Published: (2025)
by: Ngoiya, Eric, et al.
Published: (2025)
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
by: Zhu, Jiace, et al.
Published: (2026)
by: Zhu, Jiace, et al.
Published: (2026)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
EHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
by: Wang, Haochuan, et al.
Published: (2024)
by: Wang, Haochuan, et al.
Published: (2024)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
by: Fabbri, Alexander R., et al.
Published: (2025)
by: Fabbri, Alexander R., et al.
Published: (2025)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
by: Shao, Jiaqi, et al.
Published: (2025)
by: Shao, Jiaqi, et al.
Published: (2025)
INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
by: Tang, Bintao, et al.
Published: (2025)
by: Tang, Bintao, et al.
Published: (2025)
Design and Report Benchmarks for Knowledge Work
by: Hua, Yining, et al.
Published: (2026)
by: Hua, Yining, et al.
Published: (2026)
PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
by: Zhang, Jiahui, et al.
Published: (2026)
by: Zhang, Jiahui, et al.
Published: (2026)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
by: Luo, Shixian, et al.
Published: (2025)
by: Luo, Shixian, et al.
Published: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
PFMBench: Protein Foundation Model Benchmark
by: Gao, Zhangyang, et al.
Published: (2025)
by: Gao, Zhangyang, et al.
Published: (2025)
VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
by: Gu, Zixuan, et al.
Published: (2025)
by: Gu, Zixuan, et al.
Published: (2025)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
Benchmarking LLMs for Predictive Applications in the Intensive Care Units
by: Malhotra, Chehak, et al.
Published: (2025)
by: Malhotra, Chehak, et al.
Published: (2025)
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
by: Zhao, Chenchen, et al.
Published: (2025)
by: Zhao, Chenchen, et al.
Published: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
by: Wu, Yue, et al.
Published: (2023)
by: Wu, Yue, et al.
Published: (2023)
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
by: Zhu, Yuke, et al.
Published: (2020)
by: Zhu, Yuke, et al.
Published: (2020)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Similar Items
-
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026) -
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
by: Akhtar, Mubashara, et al.
Published: (2026) -
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026) -
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
by: Xu, Chengliang, et al.
Published: (2026) -
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)