ReEfBench: Quantifying the Reasoning Efficiency of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Zhizhang, Gu, Yuancheng, Hu, Chenkai, Liu, Hanmeng, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
by: Liu, Hanmeng, et al.
Published: (2025)
by: Liu, Hanmeng, et al.
Published: (2025)
Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
by: FU, Zhizhang, et al.
Published: (2025)
by: FU, Zhizhang, et al.
Published: (2025)
Break the Chain: Large Language Models Can be Shortcut Reasoners
by: Ding, Mengru, et al.
Published: (2024)
by: Ding, Mengru, et al.
Published: (2024)
Logical Reasoning in Large Language Models: A Survey
by: Liu, Hanmeng, et al.
Published: (2025)
by: Liu, Hanmeng, et al.
Published: (2025)
Logic Agent: Enhancing Validity with Logic Rule Invocation
by: Liu, Hanmeng, et al.
Published: (2024)
by: Liu, Hanmeng, et al.
Published: (2024)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
by: Ramezanali, Mohammad, et al.
Published: (2025)
by: Ramezanali, Mohammad, et al.
Published: (2025)
GLoRE: Evaluating Logical Reasoning of Large Language Models
by: liu, Hanmeng, et al.
Published: (2023)
by: liu, Hanmeng, et al.
Published: (2023)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
OckBench: Measuring the Efficiency of LLM Reasoning
by: Du, Zheng, et al.
Published: (2025)
by: Du, Zheng, et al.
Published: (2025)
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
by: Kirichenko, Polina, et al.
Published: (2025)
by: Kirichenko, Polina, et al.
Published: (2025)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
by: Zhang, Ziqian, et al.
Published: (2026)
by: Zhang, Ziqian, et al.
Published: (2026)
Know Your Needs Better: Towards Structured Understanding of Marketer Demands with Analogical Reasoning Augmented LLMs
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
by: Luo, Shixian, et al.
Published: (2025)
by: Luo, Shixian, et al.
Published: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
by: Weng, Shichao, et al.
Published: (2025)
by: Weng, Shichao, et al.
Published: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
by: Yao, Jianzhu, et al.
Published: (2025)
by: Yao, Jianzhu, et al.
Published: (2025)
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
by: Wang, Xukai, et al.
Published: (2025)
by: Wang, Xukai, et al.
Published: (2025)
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
by: Li, Kuan, et al.
Published: (2026)
by: Li, Kuan, et al.
Published: (2026)
How Reliable are LLMs for Reasoning on the Re-ranking task?
by: Islam, Nafis Tanveer, et al.
Published: (2025)
by: Islam, Nafis Tanveer, et al.
Published: (2025)
Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models
by: Wang, Yizhi, et al.
Published: (2026)
by: Wang, Yizhi, et al.
Published: (2026)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
by: Xie, Danning, et al.
Published: (2025)
by: Xie, Danning, et al.
Published: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
by: Xu, Ruiling, et al.
Published: (2025)
by: Xu, Ruiling, et al.
Published: (2025)
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
by: Li, Xuzhao, et al.
Published: (2025)
by: Li, Xuzhao, et al.
Published: (2025)
Large Language Model's Multi-Capability Alignment in Biomedical Domain
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
by: Zhang, Chenkai, et al.
Published: (2025)
by: Zhang, Chenkai, et al.
Published: (2025)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
by: Chen, Mingyang, et al.
Published: (2025)
by: Chen, Mingyang, et al.
Published: (2025)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
by: Lou, Siyu, et al.
Published: (2024)
by: Lou, Siyu, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
by: Shi, Shaojie, et al.
Published: (2026)
by: Shi, Shaojie, et al.
Published: (2026)
Similar Items
-
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
by: Liu, Hanmeng, et al.
Published: (2025) -
Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
by: FU, Zhizhang, et al.
Published: (2025) -
Break the Chain: Large Language Models Can be Shortcut Reasoners
by: Ding, Mengru, et al.
Published: (2024) -
Logical Reasoning in Large Language Models: A Survey
by: Liu, Hanmeng, et al.
Published: (2025) -
Logic Agent: Enhancing Validity with Logic Rule Invocation
by: Liu, Hanmeng, et al.
Published: (2024)