Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adib, Shefayat E Shams, Sani, Ahmed Alfey, Esham, Ekramul Alam, Abrar, Ajwad, Chowdhury, Tareque Mohmud |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
von: Sani, Ahmed Alfey, et al.
Veröffentlicht: (2026)
von: Sani, Ahmed Alfey, et al.
Veröffentlicht: (2026)
BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering
von: Sultana, Sadia, et al.
Veröffentlicht: (2025)
von: Sultana, Sadia, et al.
Veröffentlicht: (2025)
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
von: Rafid, Ahmed, et al.
Veröffentlicht: (2026)
von: Rafid, Ahmed, et al.
Veröffentlicht: (2026)
From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
von: Sakib, Shadman, et al.
Veröffentlicht: (2025)
von: Sakib, Shadman, et al.
Veröffentlicht: (2025)
Nuclei Instance Segmentation of Cryosectioned H&E Stained Histological Images using Triple U-Net Architecture
von: Ahmed, Zarif, et al.
Veröffentlicht: (2024)
von: Ahmed, Zarif, et al.
Veröffentlicht: (2024)
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
von: Alam, Kazi Samin Yasar, et al.
Veröffentlicht: (2026)
von: Alam, Kazi Samin Yasar, et al.
Veröffentlicht: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings
von: Muna, Saiyma Sittul, et al.
Veröffentlicht: (2025)
von: Muna, Saiyma Sittul, et al.
Veröffentlicht: (2025)
Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
von: Li, Junlong, et al.
Veröffentlicht: (2022)
von: Li, Junlong, et al.
Veröffentlicht: (2022)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
von: Eyasir, Abrar, et al.
Veröffentlicht: (2026)
von: Eyasir, Abrar, et al.
Veröffentlicht: (2026)
A Pan-cancer Classification Model using Multi-view Feature Selection Method and Ensemble Classifier
von: Chowdhury, Tareque Mohmud, et al.
Veröffentlicht: (2025)
von: Chowdhury, Tareque Mohmud, et al.
Veröffentlicht: (2025)
AREG: Adversarial Resource Extraction Game for Evaluating Persuasion and Resistance in Large Language Models
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
von: Muhamed, Aashiq
Veröffentlicht: (2025)
von: Muhamed, Aashiq
Veröffentlicht: (2025)
DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models
von: Rawat, Rajat, et al.
Veröffentlicht: (2024)
von: Rawat, Rajat, et al.
Veröffentlicht: (2024)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations
von: Mosharrof, Adib, et al.
Veröffentlicht: (2024)
von: Mosharrof, Adib, et al.
Veröffentlicht: (2024)
Precision Cancer Classification and Biomarker Identification from mRNA Gene Expression via Dimensionality Reduction and Explainable AI
von: Tabassum, Farzana, et al.
Veröffentlicht: (2024)
von: Tabassum, Farzana, et al.
Veröffentlicht: (2024)
Social media polarization during conflict: Insights from an ideological stance dataset on Israel-Palestine Reddit comments
von: Ali, Hasin Jawad, et al.
Veröffentlicht: (2025)
von: Ali, Hasin Jawad, et al.
Veröffentlicht: (2025)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
von: Arpa, Zaara Zabeen, et al.
Veröffentlicht: (2025)
von: Arpa, Zaara Zabeen, et al.
Veröffentlicht: (2025)
Zero-Shot Keyphrase Generation: Investigating Specialized Instructions and Multi-Sample Aggregation on Large Language Models
von: Mohan, Jayanth, et al.
Veröffentlicht: (2025)
von: Mohan, Jayanth, et al.
Veröffentlicht: (2025)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
Zero-Shot Generalizable End-to-End Task-Oriented Dialog System using Context Summarization and Domain Schema
von: Mosharrof, Adib, et al.
Veröffentlicht: (2023)
von: Mosharrof, Adib, et al.
Veröffentlicht: (2023)
Zero-Shot Verification-guided Chain of Thoughts
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2025)
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2025)
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
von: Chowdhury, Mohammed Nowshad Ruhani, et al.
Veröffentlicht: (2026)
von: Chowdhury, Mohammed Nowshad Ruhani, et al.
Veröffentlicht: (2026)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
von: Hong, Chang, et al.
Veröffentlicht: (2025)
von: Hong, Chang, et al.
Veröffentlicht: (2025)
Evaluating Zero-Shot Long-Context LLM Compression
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
An Empirical Evaluation of Large Language Models on Consumer Health Questions
von: Abrar, Moaiz, et al.
Veröffentlicht: (2024)
von: Abrar, Moaiz, et al.
Veröffentlicht: (2024)
CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
von: Ali, Hasin Jawad, et al.
Veröffentlicht: (2025)
von: Ali, Hasin Jawad, et al.
Veröffentlicht: (2025)
Evaluating Zero-Shot Multilingual Aspect-Based Sentiment Analysis with Large Language Models
von: Wu, Chengyan, et al.
Veröffentlicht: (2024)
von: Wu, Chengyan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026) -
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026) -
Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
von: Sani, Ahmed Alfey, et al.
Veröffentlicht: (2026) -
BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering
von: Sultana, Sadia, et al.
Veröffentlicht: (2025) -
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
von: Abrar, Ajwad, et al.
Veröffentlicht: (2025)