Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
Fuente:
arXiv
Saved in:
| Main Authors: | Sahu, Gaurav, Charlin, Laurent, Pal, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
TEARS: Textual Representations for Scrutable Recommendations
by: Penaloza, Emiliano, et al.
Published: (2024)
by: Penaloza, Emiliano, et al.
Published: (2024)
Policy-Grounded Dynamic Facet Suggestions for Job Search
by: Xu, Dan, et al.
Published: (2026)
by: Xu, Dan, et al.
Published: (2026)
Reading Between the Citations: A Typed Claim Network for Scientific Literature
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
Benchmark for Evaluation and Analysis of Citation Recommendation Models
by: Maharjan, Puja
Published: (2024)
by: Maharjan, Puja
Published: (2024)
Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
by: Cattaneo, Alberto, et al.
Published: (2025)
by: Cattaneo, Alberto, et al.
Published: (2025)
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
by: Zhang, Dake, et al.
Published: (2026)
by: Zhang, Dake, et al.
Published: (2026)
KG-First, LLM-Fallback: A Hybrid Microservice for Grounded Skill Search and Explanation
by: Le, Ngoc Luyen, et al.
Published: (2026)
by: Le, Ngoc Luyen, et al.
Published: (2026)
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
by: Xu, Yumo, et al.
Published: (2025)
by: Xu, Yumo, et al.
Published: (2025)
Plans for Evaluating Structured Generative Search Summaries
by: Sakai, Tetsuya, et al.
Published: (2026)
by: Sakai, Tetsuya, et al.
Published: (2026)
Document Attribution: Examining Citation Relationships using Large Language Models
by: Rawte, Vipula, et al.
Published: (2025)
by: Rawte, Vipula, et al.
Published: (2025)
Introducing ORKG ASK: an AI-driven Scholarly Literature Search and Exploration System Taking a Neuro-Symbolic Approach
by: Oelen, Allard, et al.
Published: (2025)
by: Oelen, Allard, et al.
Published: (2025)
CiteBART: Learning to Generate Citations for Local Citation Recommendation
by: Çelik, Ege Yiğit, et al.
Published: (2024)
by: Çelik, Ege Yiğit, et al.
Published: (2024)
Extracting Research Instruments from Educational Literature Using LLMs
by: Yoo, Jiseung, et al.
Published: (2025)
by: Yoo, Jiseung, et al.
Published: (2025)
Deep Learning-Based Approach for Improving Relational Aggregated Search
by: Soliman, Sara Saad, et al.
Published: (2025)
by: Soliman, Sara Saad, et al.
Published: (2025)
A Conceptual Framework for Conversational Search and Recommendation: Conceptualizing Agent-Human Interactions During the Conversational Search Process
by: Azzopardi, Leif, et al.
Published: (2024)
by: Azzopardi, Leif, et al.
Published: (2024)
DeepShop: A Benchmark for Deep Research Shopping Agents
by: Lyu, Yougang, et al.
Published: (2025)
by: Lyu, Yougang, et al.
Published: (2025)
Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility
by: Liu, Zikang, et al.
Published: (2026)
by: Liu, Zikang, et al.
Published: (2026)
Citation-Driven Multi-View Training for Patent Embeddings: QaECTER and Sophia-Bench
by: Djemmal, Younes, et al.
Published: (2026)
by: Djemmal, Younes, et al.
Published: (2026)
FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations
by: Peng, Yixing, et al.
Published: (2025)
by: Peng, Yixing, et al.
Published: (2025)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
by: Hsu, Tz-Huan, et al.
Published: (2026)
by: Hsu, Tz-Huan, et al.
Published: (2026)
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
by: Sahu, Gaurav, et al.
Published: (2025)
by: Sahu, Gaurav, et al.
Published: (2025)
AI Prior Art Search: Semantic Clusters and Evaluation Infrastructure
by: Genin, Boris, et al.
Published: (2025)
by: Genin, Boris, et al.
Published: (2025)
Evaluating Embedding Models and Pipeline Optimization for AI Search Quality
by: Zhong, Philip, et al.
Published: (2025)
by: Zhong, Philip, et al.
Published: (2025)
Evaluating Search Engines and Large Language Models for Answering Health Questions
by: Fernández-Pichel, Marcos, et al.
Published: (2024)
by: Fernández-Pichel, Marcos, et al.
Published: (2024)
Towards Personalized Deep Research: Benchmarks and Evaluations
by: Liang, Yuan, et al.
Published: (2025)
by: Liang, Yuan, et al.
Published: (2025)
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented Generation
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
Modernizing Facebook Scoped Search: Keyword and Embedding Hybrid Retrieval with LLM Evaluation
by: Su, Yongye, et al.
Published: (2025)
by: Su, Yongye, et al.
Published: (2025)
Evaluating AI Recruitment Sourcing Tools by Human Preference
by: Slaykovskiy, Vladimir, et al.
Published: (2025)
by: Slaykovskiy, Vladimir, et al.
Published: (2025)
Self-Optimizing Multi-Agent Systems for Deep Research
by: Câmara, Arthur, et al.
Published: (2026)
by: Câmara, Arthur, et al.
Published: (2026)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
ResearchPilot: A Local-First Multi-Agent System for Literature Synthesis and Related Work Drafting
by: Zhang, Peng
Published: (2026)
by: Zhang, Peng
Published: (2026)
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
by: Seo, Yongsik, et al.
Published: (2026)
by: Seo, Yongsik, et al.
Published: (2026)
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
by: Zhang, Haoxiang, et al.
Published: (2026)
by: Zhang, Haoxiang, et al.
Published: (2026)
Similar Items
-
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024) -
TEARS: Textual Representations for Scrutable Recommendations
by: Penaloza, Emiliano, et al.
Published: (2024) -
Policy-Grounded Dynamic Facet Suggestions for Job Search
by: Xu, Dan, et al.
Published: (2026) -
Reading Between the Citations: A Typed Claim Network for Scientific Literature
by: Ding, Ning, et al.
Published: (2026) -
Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval
by: Mishra, Suyash, et al.
Published: (2026)