OckBench: Measuring the Efficiency of LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Zheng, Kang, Hao, Han, Song, Krishna, Tushar, Zhu, Ligeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Slm-mux: Orchestrating small language models for reasoning
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
by: Yu, Chengjun, et al.
Published: (2026)
by: Yu, Chengjun, et al.
Published: (2026)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
by: Lin, Leon, et al.
Published: (2025)
by: Lin, Leon, et al.
Published: (2025)
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
by: Schmidt, Jan-Philipp
Published: (2026)
by: Schmidt, Jan-Philipp
Published: (2026)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
by: Jing, Huihao, et al.
Published: (2025)
by: Jing, Huihao, et al.
Published: (2025)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
by: Chen, Xuhang, et al.
Published: (2025)
by: Chen, Xuhang, et al.
Published: (2025)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
by: Zhao, Junjie, et al.
Published: (2026)
by: Zhao, Junjie, et al.
Published: (2026)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
by: Mittal, Chinmay, et al.
Published: (2024)
by: Mittal, Chinmay, et al.
Published: (2024)
WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives
by: Yu, Yongan, et al.
Published: (2025)
by: Yu, Yongan, et al.
Published: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
by: Markowitz, Elan, et al.
Published: (2025)
by: Markowitz, Elan, et al.
Published: (2025)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
by: Chu, Zheng, et al.
Published: (2023)
by: Chu, Zheng, et al.
Published: (2023)
On Path to Multimodal Historical Reasoning: HistBench and HistAgent
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models
by: Jiang, Kehan, et al.
Published: (2026)
by: Jiang, Kehan, et al.
Published: (2026)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
by: Tosato, Tommaso, et al.
Published: (2025)
by: Tosato, Tommaso, et al.
Published: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples
by: Yu, Fangxu, et al.
Published: (2024)
by: Yu, Fangxu, et al.
Published: (2024)
Adaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, and Graph Structures
by: Pandey, Tushar, et al.
Published: (2025)
by: Pandey, Tushar, et al.
Published: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
by: Wu, Junde, et al.
Published: (2025)
by: Wu, Junde, et al.
Published: (2025)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
TAP4LLM: Table Provider on Sampling, Augmenting, and Packing Semi-structured Data for Large Language Model Reasoning
by: Sui, Yuan, et al.
Published: (2023)
by: Sui, Yuan, et al.
Published: (2023)
AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
by: Zhang, Xiao Yu Cindy, et al.
Published: (2025)
by: Zhang, Xiao Yu Cindy, et al.
Published: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Similar Items
-
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025) -
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
by: Kang, Hao, et al.
Published: (2024) -
Slm-mux: Orchestrating small language models for reasoning
by: Wang, Chenyu, et al.
Published: (2025) -
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024) -
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
by: Yu, Chengjun, et al.
Published: (2026)