HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Zhuohang, Wu, Pangjing, Liang, Ziran, Chen, Peter Q., Yuan, Xu, Jia, Ye, Tu, Jiancheng, Li, Chen, Ng, Peter H. F., Li, Qing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
di: Xu, Zelin, et al.
Pubblicazione: (2026)
di: Xu, Zelin, et al.
Pubblicazione: (2026)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
di: Zhou, Yu, et al.
Pubblicazione: (2024)
di: Zhou, Yu, et al.
Pubblicazione: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025)
di: Lin, Jingru, et al.
Pubblicazione: (2025)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
di: Miao, Tingjia, et al.
Pubblicazione: (2026)
di: Miao, Tingjia, et al.
Pubblicazione: (2026)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
di: Ma, Haokai, et al.
Pubblicazione: (2025)
di: Ma, Haokai, et al.
Pubblicazione: (2025)
HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
di: Zhang, Yaping, et al.
Pubblicazione: (2025)
di: Zhang, Yaping, et al.
Pubblicazione: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
di: Hong, Mengze, et al.
Pubblicazione: (2025)
di: Hong, Mengze, et al.
Pubblicazione: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
di: Wang, Qinsi, et al.
Pubblicazione: (2026)
di: Wang, Qinsi, et al.
Pubblicazione: (2026)
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
di: Jiang, Dingfeng, et al.
Pubblicazione: (2026)
di: Jiang, Dingfeng, et al.
Pubblicazione: (2026)
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
di: Kachroo, Darsh, et al.
Pubblicazione: (2026)
di: Kachroo, Darsh, et al.
Pubblicazione: (2026)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
di: Yang, Jialin, et al.
Pubblicazione: (2025)
di: Yang, Jialin, et al.
Pubblicazione: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
di: Zhu, Hengchuan, et al.
Pubblicazione: (2025)
di: Zhu, Hengchuan, et al.
Pubblicazione: (2025)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
di: Lu, Yuyin, et al.
Pubblicazione: (2026)
di: Lu, Yuyin, et al.
Pubblicazione: (2026)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
di: Liu, Junqi, et al.
Pubblicazione: (2025)
di: Liu, Junqi, et al.
Pubblicazione: (2025)
ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs
di: Fang, Rui, et al.
Pubblicazione: (2026)
di: Fang, Rui, et al.
Pubblicazione: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
di: Guo, Zichun, et al.
Pubblicazione: (2026)
di: Guo, Zichun, et al.
Pubblicazione: (2026)
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
di: Hashemi, Masoud, et al.
Pubblicazione: (2025)
di: Hashemi, Masoud, et al.
Pubblicazione: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
di: Wang, Ben, et al.
Pubblicazione: (2026)
di: Wang, Ben, et al.
Pubblicazione: (2026)
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
di: Li, Hongxiang, et al.
Pubblicazione: (2025)
di: Li, Hongxiang, et al.
Pubblicazione: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
di: Cheng, Junhang, et al.
Pubblicazione: (2026)
di: Cheng, Junhang, et al.
Pubblicazione: (2026)
SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses
di: Jiang, Zhuohang, et al.
Pubblicazione: (2026)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2026)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
di: Ji, An-Yang, et al.
Pubblicazione: (2026)
di: Ji, An-Yang, et al.
Pubblicazione: (2026)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
di: Luo, Shixian, et al.
Pubblicazione: (2025)
di: Luo, Shixian, et al.
Pubblicazione: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
HiLight: A Hierarchy-aware Light Global Model with Hierarchical Local ConTrastive Learning
di: Chen, Zhijian, et al.
Pubblicazione: (2024)
di: Chen, Zhijian, et al.
Pubblicazione: (2024)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and Challenges
di: Quan, Pengrui, et al.
Pubblicazione: (2025)
di: Quan, Pengrui, et al.
Pubblicazione: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
di: Chen, Qirui, et al.
Pubblicazione: (2026)
di: Chen, Qirui, et al.
Pubblicazione: (2026)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
di: Li, Haoming, et al.
Pubblicazione: (2025)
di: Li, Haoming, et al.
Pubblicazione: (2025)
Hi-GMAE: Hierarchical Graph Masked Autoencoders
di: Liu, Chuang, et al.
Pubblicazione: (2024)
di: Liu, Chuang, et al.
Pubblicazione: (2024)
BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications
di: Hao, Jianing, et al.
Pubblicazione: (2026)
di: Hao, Jianing, et al.
Pubblicazione: (2026)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
di: Srinivasa, Rakshith S, et al.
Pubblicazione: (2025)
di: Srinivasa, Rakshith S, et al.
Pubblicazione: (2025)
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
di: Hu, Zhenyu, et al.
Pubblicazione: (2026)
di: Hu, Zhenyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025) -
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
di: Xu, Zelin, et al.
Pubblicazione: (2026) -
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
di: Zhou, Yu, et al.
Pubblicazione: (2024) -
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025) -
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
di: Miao, Tingjia, et al.
Pubblicazione: (2026)