CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Zicheng, Gou, Zhibin, Liang, Tian, Luo, Ruilin, Liu, Haowei, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
Progressive Knowledge Graph Completion
by: Li, Jiayi, et al.
Published: (2024)
by: Li, Jiayi, et al.
Published: (2024)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
DeepCritic: Deliberate Critique with Large Language Models
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
by: Tang, Zhengyang, et al.
Published: (2025)
by: Tang, Zhengyang, et al.
Published: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
by: Ramezanali, Mohammad, et al.
Published: (2025)
by: Ramezanali, Mohammad, et al.
Published: (2025)
Generative Evaluation of Complex Reasoning in Large Language Models
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
ProcBench: Benchmark for Multi-Step Reasoning and Following Procedure
by: Fujisawa, Ippei, et al.
Published: (2024)
by: Fujisawa, Ippei, et al.
Published: (2024)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
by: Ji, An-Yang, et al.
Published: (2026)
by: Ji, An-Yang, et al.
Published: (2026)
Self-Evolving Critique Abilities in Large Language Models
by: Tang, Zhengyang, et al.
Published: (2025)
by: Tang, Zhengyang, et al.
Published: (2025)
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
by: Gu, Tianle, et al.
Published: (2026)
by: Gu, Tianle, et al.
Published: (2026)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
by: Zhao, Zehua, et al.
Published: (2025)
by: Zhao, Zehua, et al.
Published: (2025)
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
by: Ye, Tian, et al.
Published: (2024)
by: Ye, Tian, et al.
Published: (2024)
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
by: Zhang, Jinghao, et al.
Published: (2025)
by: Zhang, Jinghao, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
by: He, Linyang, et al.
Published: (2026)
by: He, Linyang, et al.
Published: (2026)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
VL-RouterBench: A Benchmark for Vision-Language Model Routing
by: Huang, Zhehao, et al.
Published: (2025)
by: Huang, Zhehao, et al.
Published: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
by: Sim, Shamus, et al.
Published: (2024)
by: Sim, Shamus, et al.
Published: (2024)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
by: Yuan, Jiarui, et al.
Published: (2026)
by: Yuan, Jiarui, et al.
Published: (2026)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
by: Li, Junsong, et al.
Published: (2025)
by: Li, Junsong, et al.
Published: (2025)
From Faithfulness to Correctness: Generative Reward Models that Think Critically
by: Ma, Qiyao, et al.
Published: (2025)
by: Ma, Qiyao, et al.
Published: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
by: Iwase, Naoto, et al.
Published: (2025)
by: Iwase, Naoto, et al.
Published: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
DataSciBench: An LLM Agent Benchmark for Data Science
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
by: Xu, Yinggan, et al.
Published: (2025)
by: Xu, Yinggan, et al.
Published: (2025)
Similar Items
-
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024) -
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024) -
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025) -
Progressive Knowledge Graph Completion
by: Li, Jiayi, et al.
Published: (2024) -
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)