Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shan, Li, Yingya, Lu, Sheng, Van, Hoang, Aerts, Hugo JWL, Savova, Guergana K., Bitterman, Danielle S. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information
by: Li, Yingya, et al.
Published: (2024)
by: Li, Yingya, et al.
Published: (2024)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
by: Guevara, Marco, et al.
Published: (2023)
by: Guevara, Marco, et al.
Published: (2023)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models
by: Bui, Nghia, et al.
Published: (2025)
by: Bui, Nghia, et al.
Published: (2025)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
by: Gallifant, Jack, et al.
Published: (2024)
by: Gallifant, Jack, et al.
Published: (2024)
Gradable ChatGPT Translation Evaluation
by: Jiao, Hui, et al.
Published: (2024)
by: Jiao, Hui, et al.
Published: (2024)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
AI and the Law: Evaluating ChatGPT's Performance in Legal Classification
by: Weichbroth, Pawel
Published: (2025)
by: Weichbroth, Pawel
Published: (2025)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Artificial Intelligence in the Legal Field: Law Students Perspective
by: Andreeva, Daniela, et al.
Published: (2024)
by: Andreeva, Daniela, et al.
Published: (2024)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
Large language models require a new form of oversight: capability-based monitoring
by: Kellogg, Katherine C., et al.
Published: (2025)
by: Kellogg, Katherine C., et al.
Published: (2025)
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
by: Yang, Lingyi, et al.
Published: (2023)
by: Yang, Lingyi, et al.
Published: (2023)
A Survey on the Real Power of ChatGPT
by: Liu, Ming, et al.
Published: (2024)
by: Liu, Ming, et al.
Published: (2024)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
by: Tu, Shangqing, et al.
Published: (2023)
by: Tu, Shangqing, et al.
Published: (2023)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
Evaluating ChatGPT on Nuclear Domain-Specific Data
by: Anwar, Muhammad, et al.
Published: (2024)
by: Anwar, Muhammad, et al.
Published: (2024)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025)
by: Chen, Shan, et al.
Published: (2025)
ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT
by: Wei, Xiang, et al.
Published: (2023)
by: Wei, Xiang, et al.
Published: (2023)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
by: Mao, Rui, et al.
Published: (2023)
by: Mao, Rui, et al.
Published: (2023)
Generating Hard-Negative Out-of-Scope Data with ChatGPT for Intent Classification
by: Li, Zhijian, et al.
Published: (2024)
by: Li, Zhijian, et al.
Published: (2024)
Exploring ChatGPT's Capabilities on Vulnerability Management
by: Liu, Peiyu, et al.
Published: (2023)
by: Liu, Peiyu, et al.
Published: (2023)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Emergence of a phonological bias in ChatGPT
by: Toro, Juan Manuel
Published: (2023)
by: Toro, Juan Manuel
Published: (2023)
A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
by: Green, Breanna E., et al.
Published: (2025)
by: Green, Breanna E., et al.
Published: (2025)
Developing ChatGPT for Biology and Medicine: A Complete Review of Biomedical Question Answering
by: Li, Qing, et al.
Published: (2024)
by: Li, Qing, et al.
Published: (2024)
Adapting Abstract Meaning Representation Parsing to the Clinical Narrative -- the SPRING THYME parser
by: Cai, Jon Z., et al.
Published: (2024)
by: Cai, Jon Z., et al.
Published: (2024)
ChatGPT as speechwriter for the French presidents
by: Labbé, Dominique, et al.
Published: (2024)
by: Labbé, Dominique, et al.
Published: (2024)
WildChat: 1M ChatGPT Interaction Logs in the Wild
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Is ChatGPT the Future of Causal Text Mining? A Comprehensive Evaluation and Analysis
by: Takayanagi, Takehiro, et al.
Published: (2024)
by: Takayanagi, Takehiro, et al.
Published: (2024)
Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability
by: Chen, Wei-Rui, et al.
Published: (2023)
by: Chen, Wei-Rui, et al.
Published: (2023)
The Human and the Mechanical: logos, truthfulness, and ChatGPT
by: Giannakidou, Anastasia, et al.
Published: (2024)
by: Giannakidou, Anastasia, et al.
Published: (2024)
Does ChatGPT Have a Mind?
by: Goldstein, Simon, et al.
Published: (2024)
by: Goldstein, Simon, et al.
Published: (2024)
Can ChatGPT Learn to Count Letters?
by: Conde, Javier, et al.
Published: (2025)
by: Conde, Javier, et al.
Published: (2025)
On Repairing Quantum Programs Using ChatGPT
by: Guo, Xiaoyu, et al.
Published: (2024)
by: Guo, Xiaoyu, et al.
Published: (2024)
Similar Items
-
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information
by: Li, Yingya, et al.
Published: (2024) -
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
by: Guevara, Marco, et al.
Published: (2023) -
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025) -
Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models
by: Bui, Nghia, et al.
Published: (2025) -
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)