PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Daoyu, Cheng, Mingyue, Yu, Shuo, Liu, Zirui, Guo, Ze, Li, Xin, Liu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization
by: Pan, Tingyue, et al.
Published: (2026)
by: Pan, Tingyue, et al.
Published: (2026)
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
by: Zhou, Yitong, et al.
Published: (2026)
by: Zhou, Yitong, et al.
Published: (2026)
Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study
by: Yu, Shuo, et al.
Published: (2024)
by: Yu, Shuo, et al.
Published: (2024)
am-ELO: A Stable Framework for Arena-based LLM Evaluation
by: Liu, Zirui, et al.
Published: (2025)
by: Liu, Zirui, et al.
Published: (2025)
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
by: Tao, Xiaoyu, et al.
Published: (2026)
by: Tao, Xiaoyu, et al.
Published: (2026)
TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning
by: Jiang, Chuang, et al.
Published: (2025)
by: Jiang, Chuang, et al.
Published: (2025)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
A Survey on Knowledge-Oriented Retrieval-Augmented Generation
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
Revisiting the Solution of Meta KDD Cup 2024: CRAG
by: Ouyang, Jie, et al.
Published: (2024)
by: Ouyang, Jie, et al.
Published: (2024)
SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation
by: Zhao, Jiahao, et al.
Published: (2026)
by: Zhao, Jiahao, et al.
Published: (2026)
Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs
by: Zhou, Yitong, et al.
Published: (2025)
by: Zhou, Yitong, et al.
Published: (2025)
MemWeaver: A Hierarchical Memory from Textual Interactive Behaviors for Personalized Generation
by: Yu, Shuo, et al.
Published: (2025)
by: Yu, Shuo, et al.
Published: (2025)
PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
by: Xiong, Lei, et al.
Published: (2026)
by: Xiong, Lei, et al.
Published: (2026)
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation
by: Ouyang, Jie, et al.
Published: (2025)
by: Ouyang, Jie, et al.
Published: (2025)
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
by: Ma, Bo, et al.
Published: (2025)
by: Ma, Bo, et al.
Published: (2025)
TableTime: Reformulating Time Series Classification as Training-Free Table Understanding with Large Language Models
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
AlphaCast: A Human Wisdom-LLM Intelligence Co-Reasoning Framework for Interactive Time Series Forecasting
by: Zhang, Xiaohan, et al.
Published: (2025)
by: Zhang, Xiaohan, et al.
Published: (2025)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
CastFlow: Learning Role-Specialized Agentic Workflows for Time Series Forecasting
by: Pan, Bokai, et al.
Published: (2026)
by: Pan, Bokai, et al.
Published: (2026)
DeepXiv-SDK: An Agentic Data Interface for Scientific Literature
by: Qian, Hongjin, et al.
Published: (2026)
by: Qian, Hongjin, et al.
Published: (2026)
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
by: Liu, Miri, et al.
Published: (2026)
by: Liu, Miri, et al.
Published: (2026)
BibAgent: An Agentic Framework for Traceable Miscitation Detection in Scientific Literature
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
StaTS: Spectral Trajectory Schedule Learning for Adaptive Time Series Forecasting with Frequency Guided Denoiser
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts
by: Bao, Riyang, et al.
Published: (2026)
by: Bao, Riyang, et al.
Published: (2026)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
by: Wu, Junde, et al.
Published: (2025)
by: Wu, Junde, et al.
Published: (2025)
Towards Stable and Structured Time Series Generation with Perturbation-Aware Flow Matching
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
by: Zhu, Yakun, et al.
Published: (2025)
by: Zhu, Yakun, et al.
Published: (2025)
Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
by: Zhou, Yitong, et al.
Published: (2024)
by: Zhou, Yitong, et al.
Published: (2024)
GeoDecider: A Coarse-to-Fine Agentic Workflow for Explainable Lithology Classification
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025)
by: Cai, Mengzhang, et al.
Published: (2025)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
by: Liu, Tengxiao, et al.
Published: (2026)
by: Liu, Tengxiao, et al.
Published: (2026)
AnomaMind: Agentic Time Series Anomaly Detection with Tool-Augmented Reasoning
by: Tao, Xiaoyu, et al.
Published: (2026)
by: Tao, Xiaoyu, et al.
Published: (2026)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
by: Xiong, Lei, et al.
Published: (2026)
by: Xiong, Lei, et al.
Published: (2026)
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
by: Lai, Siqi, et al.
Published: (2025)
by: Lai, Siqi, et al.
Published: (2025)
Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems
by: Wang, Ling, et al.
Published: (2026)
by: Wang, Ling, et al.
Published: (2026)
Similar Items
-
PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization
by: Pan, Tingyue, et al.
Published: (2026) -
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
by: Zhou, Yitong, et al.
Published: (2026) -
Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study
by: Yu, Shuo, et al.
Published: (2024) -
am-ELO: A Stable Framework for Arena-based LLM Evaluation
by: Liu, Zirui, et al.
Published: (2025) -
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
by: Tao, Xiaoyu, et al.
Published: (2026)