Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Lin, Lin, Weihong, Wu, Jinzhu, Zhu, Yongfu, Jian, Xiaoqi, Zhao, Guangxiang, Jia, Change, Zhang, Linglin, Hu, Sai-er, Wu, Yuhan, Zhang, Xiangzheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
von: Zhu, Yongfu, et al.
Veröffentlicht: (2025)
von: Zhu, Yongfu, et al.
Veröffentlicht: (2025)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
von: Zhao, Guangxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Guangxiang, et al.
Veröffentlicht: (2025)
TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought
von: Si, Jianfeng, et al.
Veröffentlicht: (2026)
von: Si, Jianfeng, et al.
Veröffentlicht: (2026)
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
von: Zhao, Guangxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Guangxiang, et al.
Veröffentlicht: (2026)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
von: Li, Yaoming, et al.
Veröffentlicht: (2026)
von: Li, Yaoming, et al.
Veröffentlicht: (2026)
Beyond Parameter Arithmetic: Sparse Complementary Fusion for Distribution-Aware Model Merging
von: Lin, Weihong, et al.
Veröffentlicht: (2026)
von: Lin, Weihong, et al.
Veröffentlicht: (2026)
LLMs' Classification Performance is Overclaimed
von: Xu, Hanzi, et al.
Veröffentlicht: (2024)
von: Xu, Hanzi, et al.
Veröffentlicht: (2024)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation
von: Zhou, Qiji, et al.
Veröffentlicht: (2025)
von: Zhou, Qiji, et al.
Veröffentlicht: (2025)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
Ideal Registration? Segmentation is All You Need
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
Rho-1: Not All Tokens Are What You Need
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
Capabilities Ain't All You Need: Measuring Propensities in AI
von: Romero-Alvarado, Daniel, et al.
Veröffentlicht: (2026)
von: Romero-Alvarado, Daniel, et al.
Veröffentlicht: (2026)
All You Need is JSON++: Declarative Inheritance and Evaluation for JSON
von: Ukai, Hiroshi
Veröffentlicht: (2025)
von: Ukai, Hiroshi
Veröffentlicht: (2025)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
von: Pathak, Aditya, et al.
Veröffentlicht: (2025)
von: Pathak, Aditya, et al.
Veröffentlicht: (2025)
Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
von: Wang, Yang, et al.
Veröffentlicht: (2024)
von: Wang, Yang, et al.
Veröffentlicht: (2024)
Exposure Bracketing Is All You Need For A High-Quality Image
von: Zhang, Zhilu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhilu, et al.
Veröffentlicht: (2024)
Agents Are All You Need for LLM Unlearning
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
ParameterNet: Parameters Are All You Need
von: Han, Kai, et al.
Veröffentlicht: (2023)
von: Han, Kai, et al.
Veröffentlicht: (2023)
Multistep Inverse Is Not All You Need
von: Levine, Alexander, et al.
Veröffentlicht: (2024)
von: Levine, Alexander, et al.
Veröffentlicht: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
Reasoning Is All You Need for Urban Planning AI
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
Learn to Think: Bootstrapping LLM Reasoning Capability Through Graph Representation Learning
von: Gao, Hang, et al.
Veröffentlicht: (2025)
von: Gao, Hang, et al.
Veröffentlicht: (2025)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
von: Bai, Yuelin, et al.
Veröffentlicht: (2024)
von: Bai, Yuelin, et al.
Veröffentlicht: (2024)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
Are Prompts All You Need? Evaluating Prompt-Based Large Language Models (LLM)s for Software Requirements Classification
von: Binkhonain, Manal, et al.
Veröffentlicht: (2025)
von: Binkhonain, Manal, et al.
Veröffentlicht: (2025)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
von: Yashwanth, Tadisetty Sai, et al.
Veröffentlicht: (2025)
von: Yashwanth, Tadisetty Sai, et al.
Veröffentlicht: (2025)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
Emu3: Next-Token Prediction is All You Need
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
von: Zhu, Yongfu, et al.
Veröffentlicht: (2025) -
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
von: Zhao, Guangxiang, et al.
Veröffentlicht: (2025) -
TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation
von: Sun, Lin, et al.
Veröffentlicht: (2025) -
BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning
von: Sun, Lin, et al.
Veröffentlicht: (2026) -
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026)