Benchmarking Retrieval-Augmented Generation for Medicine
Fuente:
arXiv
Saved in:
| Main Authors: | Xiong, Guangzhi, Jin, Qiao, Lu, Zhiyong, Zhang, Aidong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up Questions
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
MedCite: Can Language Models Generate Verifiable Text for Medicine?
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
Benchmarking Retrieval-Augmented Generation for Chemistry
by: Zhong, Xianrui, et al.
Published: (2025)
by: Zhong, Xianrui, et al.
Published: (2025)
Retrieving Counterfactuals Improves Visual In-Context Learning
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
Supervising the search process produces reliable and generalizable information-seeking agents
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
IdeaBench: Benchmarking Large Language Models for Research Idea Generation
by: Guo, Sikun, et al.
Published: (2024)
by: Guo, Sikun, et al.
Published: (2024)
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
by: Fang, Yin, et al.
Published: (2025)
by: Fang, Yin, et al.
Published: (2025)
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators
by: Wan, Nicholas, et al.
Published: (2024)
by: Wan, Nicholas, et al.
Published: (2024)
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
by: Jin, Qiao, et al.
Published: (2023)
by: Jin, Qiao, et al.
Published: (2023)
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
by: Lu, Wensheng, et al.
Published: (2025)
by: Lu, Wensheng, et al.
Published: (2025)
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution
by: Jin, Qiao, et al.
Published: (2026)
by: Jin, Qiao, et al.
Published: (2026)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024)
by: Wang, Tevin, et al.
Published: (2024)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
by: Friel, Robert, et al.
Published: (2024)
by: Friel, Robert, et al.
Published: (2024)
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
by: Sun, Yang, et al.
Published: (2025)
by: Sun, Yang, et al.
Published: (2025)
MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
by: Khandekar, Nikhil, et al.
Published: (2024)
by: Khandekar, Nikhil, et al.
Published: (2024)
Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
by: Zheng, Shunfeng, et al.
Published: (2025)
by: Zheng, Shunfeng, et al.
Published: (2025)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
by: Park, Chanhee, et al.
Published: (2025)
by: Park, Chanhee, et al.
Published: (2025)
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
by: Mao, Qianren, et al.
Published: (2024)
by: Mao, Qianren, et al.
Published: (2024)
Predictive Prefetching for Retrieval-Augmented Generation
by: Zhang, Wuyang, et al.
Published: (2026)
by: Zhang, Wuyang, et al.
Published: (2026)
Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation
by: Yue, Zhenrui, et al.
Published: (2024)
by: Yue, Zhenrui, et al.
Published: (2024)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
by: Xiao, Yilin, et al.
Published: (2026)
by: Xiao, Yilin, et al.
Published: (2026)
ZhiFangDanTai: Fine-tuning Graph-based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula
by: Zhang, ZiXuan, et al.
Published: (2025)
by: Zhang, ZiXuan, et al.
Published: (2025)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025)
by: Katsis, Yannis, et al.
Published: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Retrieval-Augmented Generation for Large Language Models: A Survey
by: Gao, Yunfan, et al.
Published: (2023)
by: Gao, Yunfan, et al.
Published: (2023)
Enhancing Large Language Models with Domain-specific Retrieval Augment Generation: A Case Study on Long-form Consumer Health Question Answering in Ophthalmology
by: Gilson, Aidan, et al.
Published: (2024)
by: Gilson, Aidan, et al.
Published: (2024)
ContextRAG: Extraction-Free Hierarchical Graph Construction for Retrieval-Augmented Generation
by: Prosvirnin, Roman, et al.
Published: (2026)
by: Prosvirnin, Roman, et al.
Published: (2026)
Similar Items
-
Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up Questions
by: Xiong, Guangzhi, et al.
Published: (2024) -
MedCite: Can Language Models Generate Verifiable Text for Medicine?
by: Wang, Xiao, et al.
Published: (2025) -
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025) -
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
by: Xiong, Guangzhi, et al.
Published: (2026) -
Benchmarking Retrieval-Augmented Generation for Chemistry
by: Zhong, Xianrui, et al.
Published: (2025)