Reviewing Scientific Papers for Critical Problems With Reasoning LLMs: Baseline Approaches and Automatic Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Tianmai M., Abernethy, Neil F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Reference Errors in Scientific Literature with Large Language Models
by: Zhang, Tianmai M., et al.
Published: (2024)
by: Zhang, Tianmai M., et al.
Published: (2024)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
by: Li, Ruochi, et al.
Published: (2025)
by: Li, Ruochi, et al.
Published: (2025)
UW-BioNLP at ChemoTimelines 2025: Thinking, Fine-Tuning, and Dictionary-Enhanced LLM Systems for Chemotherapy Timeline Extraction
by: Zhang, Tianmai M., et al.
Published: (2025)
by: Zhang, Tianmai M., et al.
Published: (2025)
Automatic Construction of Multiple Classification Dimensions for Managing Approaches in Scientific Papers
by: Ma, Bing, et al.
Published: (2025)
by: Ma, Bing, et al.
Published: (2025)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025)
by: Xu, Zhijian, et al.
Published: (2025)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
by: Li, Alan, et al.
Published: (2025)
by: Li, Alan, et al.
Published: (2025)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers
by: Tsanda, Alena, et al.
Published: (2024)
by: Tsanda, Alena, et al.
Published: (2024)
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
by: Li, Shuaimin, et al.
Published: (2025)
by: Li, Shuaimin, et al.
Published: (2025)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
MARG: Multi-Agent Review Generation for Scientific Papers
by: D'Arcy, Mike, et al.
Published: (2024)
by: D'Arcy, Mike, et al.
Published: (2024)
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
by: Cui, Hao, et al.
Published: (2025)
by: Cui, Hao, et al.
Published: (2025)
GraphReview: Scientific Paper Evaluation via LLM-Based Graph Message Passing
by: Zheng, Pujun, et al.
Published: (2026)
by: Zheng, Pujun, et al.
Published: (2026)
Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems
by: Miner, Stephen, et al.
Published: (2024)
by: Miner, Stephen, et al.
Published: (2024)
Guardrail Baselines for Unlearning in LLMs
by: Thaker, Pratiksha, et al.
Published: (2024)
by: Thaker, Pratiksha, et al.
Published: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
by: D'Arcy, Mike, et al.
Published: (2023)
by: D'Arcy, Mike, et al.
Published: (2023)
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
by: Du, Jiangshu, et al.
Published: (2024)
by: Du, Jiangshu, et al.
Published: (2024)
Mapping the Increasing Use of LLMs in Scientific Papers
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation
by: Zeng, Qi, et al.
Published: (2023)
by: Zeng, Qi, et al.
Published: (2023)
RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval
by: Oh, Minhae, et al.
Published: (2025)
by: Oh, Minhae, et al.
Published: (2025)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
by: Dreyer, Florian, et al.
Published: (2025)
by: Dreyer, Florian, et al.
Published: (2025)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
A Critical Study of Automatic Evaluation in Sign Language Translation
by: Yazdani, Shakib, et al.
Published: (2025)
by: Yazdani, Shakib, et al.
Published: (2025)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
CCSBench: Evaluating Compositional Controllability in LLMs for Scientific Document Summarization
by: Ding, Yixi, et al.
Published: (2024)
by: Ding, Yixi, et al.
Published: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
by: Andreev, Nikita, et al.
Published: (2024)
by: Andreev, Nikita, et al.
Published: (2024)
Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
by: Goldberg, Alexander, et al.
Published: (2024)
by: Goldberg, Alexander, et al.
Published: (2024)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026)
by: Xu, Wanghan, et al.
Published: (2026)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
by: Seo, Minju, et al.
Published: (2025)
by: Seo, Minju, et al.
Published: (2025)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
by: Luo, Yuchen, et al.
Published: (2026)
by: Luo, Yuchen, et al.
Published: (2026)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
by: Dan, Nifu, et al.
Published: (2025)
by: Dan, Nifu, et al.
Published: (2025)
Similar Items
-
Detecting Reference Errors in Scientific Literature with Large Language Models
by: Zhang, Tianmai M., et al.
Published: (2024) -
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
by: Li, Ruochi, et al.
Published: (2025) -
UW-BioNLP at ChemoTimelines 2025: Thinking, Fine-Tuning, and Dictionary-Enhanced LLM Systems for Chemotherapy Timeline Extraction
by: Zhang, Tianmai M., et al.
Published: (2025) -
Automatic Construction of Multiple Classification Dimensions for Managing Approaches in Scientific Papers
by: Ma, Bing, et al.
Published: (2025) -
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025)