PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yanjun, Wei, Tianxin, Zou, Jiaru, Ning, Xuying, Bei, Yuanchen, Chen, Lingjie, Rana, Simmi, Yang, Wendy H., Tong, Hanghang, He, Jingrui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026)
by: Bei, Yuanchen, et al.
Published: (2026)
MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
by: Ning, Xuying, et al.
Published: (2026)
by: Ning, Xuying, et al.
Published: (2026)
Heterogeneous Scientific Foundation Model Collaboration
by: Li, Zihao, et al.
Published: (2026)
by: Li, Zihao, et al.
Published: (2026)
Graph4MM: Weaving Multimodal Learning with Structural Information
by: Ning, Xuying, et al.
Published: (2025)
by: Ning, Xuying, et al.
Published: (2025)
EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation
by: Li, Ting-Wei, et al.
Published: (2026)
by: Li, Ting-Wei, et al.
Published: (2026)
FeDecider: An LLM-Based Framework for Federated Cross-Domain Recommendation
by: He, Xinrui, et al.
Published: (2026)
by: He, Xinrui, et al.
Published: (2026)
Harnessing Consistency for Robust Test-Time LLM Ensemble
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
LLM-Forest: Ensemble Learning of LLMs with Graph-Augmented Prompts for Data Imputation
by: He, Xinrui, et al.
Published: (2024)
by: He, Xinrui, et al.
Published: (2024)
Saffron-1: Safety Inference Scaling
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
by: Ou, Jiefu, et al.
Published: (2025)
by: Ou, Jiefu, et al.
Published: (2025)
TRIMS: Trajectory-Ranked Instruction Masked Supervision for Diffusion Language Models
by: Chen, Lingjie, et al.
Published: (2026)
by: Chen, Lingjie, et al.
Published: (2026)
Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning
by: Bao, Wenxuan, et al.
Published: (2025)
by: Bao, Wenxuan, et al.
Published: (2025)
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
by: Du, Jiangshu, et al.
Published: (2024)
by: Du, Jiangshu, et al.
Published: (2024)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
by: Wang, Daoyu, et al.
Published: (2025)
by: Wang, Daoyu, et al.
Published: (2025)
Mapping the Increasing Use of LLMs in Scientific Papers
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
by: Xiong, Lei, et al.
Published: (2026)
by: Xiong, Lei, et al.
Published: (2026)
Subspace Alignment for Vision-Language Model Test-time Adaptation
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
PageRank Bandits for Link Prediction
by: Ban, Yikun, et al.
Published: (2024)
by: Ban, Yikun, et al.
Published: (2024)
AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
by: Cui, Chengming, et al.
Published: (2026)
by: Cui, Chengming, et al.
Published: (2026)
WAPITI: A Watermark for Finetuned Open-Source LLMs
by: Chen, Lingjie, et al.
Published: (2024)
by: Chen, Lingjie, et al.
Published: (2024)
CLIMB: Class-imbalanced Learning Benchmark on Tabular Data
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
Can LLMs Review Scientific Papers?
by: Ban, Byunghyun
Published: (2026)
by: Ban, Byunghyun
Published: (2026)
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
Flow Matching Meets Biology and Life Science: A Survey
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
TSAQA: Time Series Analysis Question And Answering Benchmark
by: Jing, Baoyu, et al.
Published: (2026)
by: Jing, Baoyu, et al.
Published: (2026)
ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
by: Lin, Xiao, et al.
Published: (2026)
by: Lin, Xiao, et al.
Published: (2026)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
by: Pang, Wei, et al.
Published: (2025)
by: Pang, Wei, et al.
Published: (2025)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
i$^2$VAE: Interest Information Augmentation with Variational Regularizers for Cross-Domain Sequential Recommendation
by: Ning, Xuying, et al.
Published: (2024)
by: Ning, Xuying, et al.
Published: (2024)
RAG over Tables: Hierarchical Memory Index, Multi-Stage Retrieval, and Benchmarking
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
Agentic Reasoning for Large Language Models
by: Wei, Tianxin, et al.
Published: (2026)
by: Wei, Tianxin, et al.
Published: (2026)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
by: Ai, Mengting, et al.
Published: (2025)
by: Ai, Mengting, et al.
Published: (2025)
dLLM: Simple Diffusion Language Modeling
by: Zhou, Zhanhui, et al.
Published: (2026)
by: Zhou, Zhanhui, et al.
Published: (2026)
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers
by: Tsanda, Alena, et al.
Published: (2024)
by: Tsanda, Alena, et al.
Published: (2024)
SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
by: Pramanick, Shraman, et al.
Published: (2024)
by: Pramanick, Shraman, et al.
Published: (2024)
Similar Items
-
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026) -
MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
by: Ning, Xuying, et al.
Published: (2026) -
Heterogeneous Scientific Foundation Model Collaboration
by: Li, Zihao, et al.
Published: (2026) -
Graph4MM: Weaving Multimodal Learning with Structural Information
by: Ning, Xuying, et al.
Published: (2025) -
EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation
by: Li, Ting-Wei, et al.
Published: (2026)