PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yutao, Liu, Xiao, Feng, Yunhao, Ding, Jiale, Ma, Xingjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
by: Guo, Ziyi, et al.
Published: (2025)
by: Guo, Ziyi, et al.
Published: (2025)
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
by: Faizullah, Abdur Rahman Bin Md, et al.
Published: (2024)
by: Faizullah, Abdur Rahman Bin Md, et al.
Published: (2024)
Learning to Ask: Conversational Product Search via Representation Learning
by: Zou, Jie, et al.
Published: (2024)
by: Zou, Jie, et al.
Published: (2024)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
by: Huang, Youcheng, et al.
Published: (2025)
by: Huang, Youcheng, et al.
Published: (2025)
Scholar Inbox: Personalized Paper Recommendations for Scientists
by: Flicke, Markus, et al.
Published: (2025)
by: Flicke, Markus, et al.
Published: (2025)
Tuning LLMs by RAG Principles: Towards LLM-native Memory
by: Wei, Jiale, et al.
Published: (2025)
by: Wei, Jiale, et al.
Published: (2025)
Enhancing Abstractive Summarization of Scientific Papers Using Structure Information
by: Bao, Tong, et al.
Published: (2025)
by: Bao, Tong, et al.
Published: (2025)
Scientific Paper Retrieval with LLM-Guided Semantic-Based Ranking
by: Zhang, Yunyi, et al.
Published: (2025)
by: Zhang, Yunyi, et al.
Published: (2025)
NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
by: Wu, Wenqing, et al.
Published: (2026)
by: Wu, Wenqing, et al.
Published: (2026)
An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
by: Ma, Linyue, et al.
Published: (2025)
by: Ma, Linyue, et al.
Published: (2025)
Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
Search-o1: Agentic Search-Enhanced Large Reasoning Models
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
Chain of Retrieval: Multi-Aspect Iterative Search Expansion and Post-Order Search Aggregation for Full Paper Retrieval
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search
by: Zhang, Erhan, et al.
Published: (2026)
by: Zhang, Erhan, et al.
Published: (2026)
Distilling Reasoning Without Knowledge: A Framework for Reliable LLMs
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
DeepRead: Document Structure-Aware Reasoning to Enhance Agentic Search
by: Li, Zhanli, et al.
Published: (2026)
by: Li, Zhanli, et al.
Published: (2026)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Reliable Evaluation Protocol for Low-Precision Retrieval
by: Yang, Kisu, et al.
Published: (2025)
by: Yang, Kisu, et al.
Published: (2025)
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
by: Xing, Tiancheng, et al.
Published: (2025)
by: Xing, Tiancheng, et al.
Published: (2025)
HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search
by: Jin, Jiajie, et al.
Published: (2025)
by: Jin, Jiajie, et al.
Published: (2025)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
EncouRAGe: Evaluating RAG Local, Fast, and Reliable
by: Strich, Jan, et al.
Published: (2025)
by: Strich, Jan, et al.
Published: (2025)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
by: Tan, Zhiwen, et al.
Published: (2025)
by: Tan, Zhiwen, et al.
Published: (2025)
Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
by: Lin, Jianghao, et al.
Published: (2025)
by: Lin, Jianghao, et al.
Published: (2025)
Towards Personalized Deep Research: Benchmarks and Evaluations
by: Liang, Yuan, et al.
Published: (2025)
by: Liang, Yuan, et al.
Published: (2025)
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
by: Zhao, Shu, et al.
Published: (2025)
by: Zhao, Shu, et al.
Published: (2025)
Reading Between the Citations: A Typed Claim Network for Scientific Literature
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
by: Hu, Xuming, et al.
Published: (2024)
by: Hu, Xuming, et al.
Published: (2024)
RAVine: Reality-Aligned Evaluation for Agentic Search
by: Xu, Yilong, et al.
Published: (2025)
by: Xu, Yilong, et al.
Published: (2025)
GISA: A Benchmark for General Information-Seeking Assistant
by: Zhu, Yutao, et al.
Published: (2026)
by: Zhu, Yutao, et al.
Published: (2026)
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
by: Song, Huatong, et al.
Published: (2025)
by: Song, Huatong, et al.
Published: (2025)
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective
by: Liu, Yunhao, et al.
Published: (2026)
by: Liu, Yunhao, et al.
Published: (2026)
SciPIP: An LLM-based Scientific Paper Idea Proposer
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Similar Items
-
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026) -
Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
by: Guo, Ziyi, et al.
Published: (2025) -
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
by: Faizullah, Abdur Rahman Bin Md, et al.
Published: (2024) -
Learning to Ask: Conversational Product Search via Representation Learning
by: Zou, Jie, et al.
Published: (2024) -
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
by: Liu, Geng, et al.
Published: (2025)