PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiong, Lei, Yuan, Huaying, Liu, Zheng, Cao, Zhao, Dou, Zhicheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
por: Wang, Daoyu, et al.
Publicado: (2025)
por: Wang, Daoyu, et al.
Publicado: (2025)
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
por: Deng, Mengjie, et al.
Publicado: (2025)
por: Deng, Mengjie, et al.
Publicado: (2025)
Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers
por: Koloveas, Paris, et al.
Publicado: (2025)
por: Koloveas, Paris, et al.
Publicado: (2025)
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
por: Zhang, Chenghao, et al.
Publicado: (2026)
por: Zhang, Chenghao, et al.
Publicado: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
por: Li, Chuhan, et al.
Publicado: (2024)
por: Li, Chuhan, et al.
Publicado: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
por: Fang, Zhicheng, et al.
Publicado: (2026)
por: Fang, Zhicheng, et al.
Publicado: (2026)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
por: Xiong, Lei, et al.
Publicado: (2026)
por: Xiong, Lei, et al.
Publicado: (2026)
Science Across Languages: Assessing LLM Multilingual Translation of Scientific Papers
por: Kleidermacher, Hannah Calzi, et al.
Publicado: (2025)
por: Kleidermacher, Hannah Calzi, et al.
Publicado: (2025)
DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning
por: Zou, Zhuoyang, et al.
Publicado: (2026)
por: Zou, Zhuoyang, et al.
Publicado: (2026)
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
por: Wan, Shu, et al.
Publicado: (2026)
por: Wan, Shu, et al.
Publicado: (2026)
PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation via Multi-Agent LLMs
por: Zhang, Zhilin, et al.
Publicado: (2025)
por: Zhang, Zhilin, et al.
Publicado: (2025)
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
por: Yu, Bihui, et al.
Publicado: (2026)
por: Yu, Bihui, et al.
Publicado: (2026)
BloClaw: An Omniscient, Multi-Modal Agentic Workspace for Next-Generation Scientific Discovery
por: Qin, Yao, et al.
Publicado: (2026)
por: Qin, Yao, et al.
Publicado: (2026)
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
por: Chen, Yelin, et al.
Publicado: (2026)
por: Chen, Yelin, et al.
Publicado: (2026)
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
por: Song, Yiwen, et al.
Publicado: (2026)
por: Song, Yiwen, et al.
Publicado: (2026)
APRES: An Agentic Paper Revision and Evaluation System
por: Zhao, Bingchen, et al.
Publicado: (2026)
por: Zhao, Bingchen, et al.
Publicado: (2026)
Preacher: Paper-to-Video Agentic System
por: Liu, Jingwei, et al.
Publicado: (2025)
por: Liu, Jingwei, et al.
Publicado: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
por: Zhu, Zeyu, et al.
Publicado: (2025)
por: Zhu, Zeyu, et al.
Publicado: (2025)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
por: Xi, Sarina, et al.
Publicado: (2025)
por: Xi, Sarina, et al.
Publicado: (2025)
Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
por: Guo, Ziyi, et al.
Publicado: (2025)
por: Guo, Ziyi, et al.
Publicado: (2025)
Discourse-Aware Scientific Paper Recommendation via QA-Style Summarization and Multi-Level Contrastive Learning
por: Wang, Shenghua, et al.
Publicado: (2025)
por: Wang, Shenghua, et al.
Publicado: (2025)
DeepXiv-SDK: An Agentic Data Interface for Scientific Literature
por: Qian, Hongjin, et al.
Publicado: (2026)
por: Qian, Hongjin, et al.
Publicado: (2026)
Mapping the Increasing Use of LLMs in Scientific Papers
por: Liang, Weixin, et al.
Publicado: (2024)
por: Liang, Weixin, et al.
Publicado: (2024)
Paper Espresso: From Paper Overload to Research Insight
por: Du, Mingzhe, et al.
Publicado: (2026)
por: Du, Mingzhe, et al.
Publicado: (2026)
Reasoning-Aware AIGC Detection via Alignment and Reinforcement
por: Wang, Zhao, et al.
Publicado: (2026)
por: Wang, Zhao, et al.
Publicado: (2026)
GLARE: Agentic Reasoning for Legal Judgment Prediction
por: Yang, Xinyu, et al.
Publicado: (2025)
por: Yang, Xinyu, et al.
Publicado: (2025)
Agentic Mixture-of-Workflows for Multi-Modal Chemical Search
por: Callahan, Tiffany J., et al.
Publicado: (2025)
por: Callahan, Tiffany J., et al.
Publicado: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
por: Wu, Yutao, et al.
Publicado: (2025)
por: Wu, Yutao, et al.
Publicado: (2025)
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes
por: Cheng, Jiehan, et al.
Publicado: (2025)
por: Cheng, Jiehan, et al.
Publicado: (2025)
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
por: Kohler, Benjamin, et al.
Publicado: (2026)
por: Kohler, Benjamin, et al.
Publicado: (2026)
Multi-Modal Experience Inspired AI Creation
por: Cao, Qian, et al.
Publicado: (2022)
por: Cao, Qian, et al.
Publicado: (2022)
LawThinker: A Deep Research Legal Agent in Dynamic Environments
por: Yang, Xinyu, et al.
Publicado: (2026)
por: Yang, Xinyu, et al.
Publicado: (2026)
A Multi-Modal Deep Learning Framework for Pan-Cancer Prognosis
por: Zhang, Binyu, et al.
Publicado: (2025)
por: Zhang, Binyu, et al.
Publicado: (2025)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
por: Burgess, James, et al.
Publicado: (2026)
por: Burgess, James, et al.
Publicado: (2026)
Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
por: Chen, Yifei, et al.
Publicado: (2025)
por: Chen, Yifei, et al.
Publicado: (2025)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
por: Xu, Ke, et al.
Publicado: (2026)
por: Xu, Ke, et al.
Publicado: (2026)
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
por: Wang, Lintao, et al.
Publicado: (2025)
por: Wang, Lintao, et al.
Publicado: (2025)
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
por: Liu, Zhenghao, et al.
Publicado: (2025)
por: Liu, Zhenghao, et al.
Publicado: (2025)
Ejemplares similares
-
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
por: Wang, Daoyu, et al.
Publicado: (2025) -
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
por: Deng, Mengjie, et al.
Publicado: (2025) -
Accelerating Scientific Discovery with Multi-Document Summarization of Impact-Ranked Papers
por: Koloveas, Paris, et al.
Publicado: (2025) -
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
por: Zhang, Chenghao, et al.
Publicado: (2026) -
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
por: Yuan, Huaying, et al.
Publicado: (2025)