Overview of TREC 2025 Biomedical Generative Retrieval (BioGen) Track

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Deepak, Demner-Fushman, Dina, Hersh, William, Bedrick, Steven, Roberts, Kirk
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911537704206336
author Gupta, Deepak
Demner-Fushman, Dina
Hersh, William
Bedrick, Steven
Roberts, Kirk
author_facet Gupta, Deepak
Demner-Fushman, Dina
Hersh, William
Bedrick, Steven
Roberts, Kirk
contents Recent advances in large language models (LLMs) have made significant progress across multiple biomedical tasks, including biomedical question answering, lay-language summarization of the biomedical literature, and clinical note summarization. These models have demonstrated strong capabilities in processing and synthesizing complex biomedical information and in generating fluent, human-like responses. Despite these advancements, hallucinations or confabulations remain key challenges when using LLMs in biomedical and other high-stakes domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly
format Preprint
id arxiv_https___arxiv_org_abs_2603_21582
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Overview of TREC 2025 Biomedical Generative Retrieval (BioGen) Track
Gupta, Deepak
Demner-Fushman, Dina
Hersh, William
Bedrick, Steven
Roberts, Kirk
Information Retrieval
Recent advances in large language models (LLMs) have made significant progress across multiple biomedical tasks, including biomedical question answering, lay-language summarization of the biomedical literature, and clinical note summarization. These models have demonstrated strong capabilities in processing and synthesizing complex biomedical information and in generating fluent, human-like responses. Despite these advancements, hallucinations or confabulations remain key challenges when using LLMs in biomedical and other high-stakes domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly
title Overview of TREC 2025 Biomedical Generative Retrieval (BioGen) Track
topic Information Retrieval
url https://arxiv.org/abs/2603.21582