Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Vatsal, Shubham, Singh, Ayush |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
by: Vatsal, Shubham, et al.
Published: (2025)
by: Vatsal, Shubham, et al.
Published: (2025)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
by: Chen, Junying, et al.
Published: (2023)
by: Chen, Junying, et al.
Published: (2023)
ArabianGPT: Native Arabic GPT-based Large Language Model
by: Koubaa, Anis, et al.
Published: (2024)
by: Koubaa, Anis, et al.
Published: (2024)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
by: Zain, Noor Ul, et al.
Published: (2025)
by: Zain, Noor Ul, et al.
Published: (2025)
Architectural Flaw Detection in Civil Engineering Using GPT-4
by: Kumar, Saket, et al.
Published: (2024)
by: Kumar, Saket, et al.
Published: (2024)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
by: Ball, Thomas, et al.
Published: (2024)
by: Ball, Thomas, et al.
Published: (2024)
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
by: Bayani, David
Published: (2023)
by: Bayani, David
Published: (2023)
Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
by: Sato, Ryoma
Published: (2026)
by: Sato, Ryoma
Published: (2026)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Evaluating GPT's Capability in Identifying Stages of Cognitive Impairment from Electronic Health Data
by: Leng, Yu, et al.
Published: (2025)
by: Leng, Yu, et al.
Published: (2025)
Universal Neurons in GPT2 Language Models
by: Gurnee, Wes, et al.
Published: (2024)
by: Gurnee, Wes, et al.
Published: (2024)
HumanEval on Latest GPT Models -- 2024
by: Li, Daniel, et al.
Published: (2024)
by: Li, Daniel, et al.
Published: (2024)
NExT-GPT: Any-to-Any Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2023)
by: Wu, Shengqiong, et al.
Published: (2023)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages
by: Mercer, Johnathan
Published: (2024)
by: Mercer, Johnathan
Published: (2024)
Does Biomedical Training Lead to Better Medical Performance?
by: Dada, Amin, et al.
Published: (2024)
by: Dada, Amin, et al.
Published: (2024)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
by: Tang, Zihao, et al.
Published: (2024)
by: Tang, Zihao, et al.
Published: (2024)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
by: Chen, Yuhao, et al.
Published: (2023)
by: Chen, Yuhao, et al.
Published: (2023)
GPT-4o Lacks Core Features of Theory of Mind
by: Muchovej, John, et al.
Published: (2026)
by: Muchovej, John, et al.
Published: (2026)
Robustness of Large Language Models to Perturbations in Text
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
CityGPT: Empowering Urban Spatial Cognition of Large Language Models
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
GPT-3 Powered Information Extraction for Building Robust Knowledge Bases
by: Choudhury, Ritabrata Roy, et al.
Published: (2024)
by: Choudhury, Ritabrata Roy, et al.
Published: (2024)
ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change
by: Thulke, David, et al.
Published: (2024)
by: Thulke, David, et al.
Published: (2024)
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Jill Watson: A Virtual Teaching Assistant powered by ChatGPT
by: Taneja, Karan, et al.
Published: (2024)
by: Taneja, Karan, et al.
Published: (2024)
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models
by: Lv, Yaojia, et al.
Published: (2024)
by: Lv, Yaojia, et al.
Published: (2024)
Similar Items
-
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024) -
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
by: Vatsal, Shubham, et al.
Published: (2025) -
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023) -
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024) -
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024)