RAG-BioQA: A Retrieval-Augmented Generation Framework for Long-Form Biomedical Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915702737207296 |
|---|---|
| author | Panchumarthi, Lovely Yeswanth Saleti, Sumalatha Gudari, Sai Prasad Negi, Atharva Budime, Praveen Raj Upadhya, Harsit |
| author_facet | Panchumarthi, Lovely Yeswanth Saleti, Sumalatha Gudari, Sai Prasad Negi, Atharva Budime, Praveen Raj Upadhya, Harsit |
| contents | The rapidly growth of biomedical literature creates challenges acquiring specific medical information. Current biomedical question-answering systems primarily focus on short-form answers, failing to provide comprehensive explanations necessary for clinical decision-making. We present RAG-BioQA, a retrieval-augmented generation framework for long-form biomedical question answering. Our system integrates BioBERT embeddings with FAISS indexing for retrieval and a LoRA fine-tuned FLAN-T5 model for answer generation. We train on 181k QA pairs from PubMedQA, MedDialog, and MedQuAD, and evaluate on a held-out PubMedQA test set. We compare four retrieval strategies: dense retrieval (FAISS), BM25, ColBERT, and MonoT5. Our results show that domain-adapted dense retrieval outperforms zero-shot neural re-rankers, with the best configuration achieving 0.24 BLEU-1 and 0.29 ROUGE-1. Fine-tuning improves BERTScore by 81\% over the base model. We release our framework to support reproducible biomedical QA research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_01612 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RAG-BioQA: A Retrieval-Augmented Generation Framework for Long-Form Biomedical Question Answering Panchumarthi, Lovely Yeswanth Saleti, Sumalatha Gudari, Sai Prasad Negi, Atharva Budime, Praveen Raj Upadhya, Harsit Computation and Language Artificial Intelligence The rapidly growth of biomedical literature creates challenges acquiring specific medical information. Current biomedical question-answering systems primarily focus on short-form answers, failing to provide comprehensive explanations necessary for clinical decision-making. We present RAG-BioQA, a retrieval-augmented generation framework for long-form biomedical question answering. Our system integrates BioBERT embeddings with FAISS indexing for retrieval and a LoRA fine-tuned FLAN-T5 model for answer generation. We train on 181k QA pairs from PubMedQA, MedDialog, and MedQuAD, and evaluate on a held-out PubMedQA test set. We compare four retrieval strategies: dense retrieval (FAISS), BM25, ColBERT, and MonoT5. Our results show that domain-adapted dense retrieval outperforms zero-shot neural re-rankers, with the best configuration achieving 0.24 BLEU-1 and 0.29 ROUGE-1. Fine-tuning improves BERTScore by 81\% over the base model. We release our framework to support reproducible biomedical QA research. |
| title | RAG-BioQA: A Retrieval-Augmented Generation Framework for Long-Form Biomedical Question Answering |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.01612 |